AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a busy software firm facing its worst week — irate customers, critical crises, and tempting shortcuts. Now, picture AI models managing this chaos, each with its own personality and style. Which one would come out on top? The answer may surprise you.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Recently, a groundbreaking live experiment put four leading frontier AI models to the test in managing a real small company thrown into turmoil. This wasn’t a scripted demo; it was a fully watchable, auditable simulation where the models faced the same challenges as human managers do — conflicting priorities, ethical dilemmas, and pressure to deliver results.

The models included GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. They each ran a virtual version of a real software business, complete with customer complaints, internal crises, and manipulative tactics. The goal? To see if AI could not only diagnose issues but also follow through, stay honest, and close deals.

Within this challenging environment, all four models successfully identified every crisis and refused every unethical manipulation attempt, including social engineering scams like fake CEO messages and staged reporter inquiries. This underlines an essential point: these models understand integrity under pressure. Notably, all models exhibited the same basic competency in crisis detection and refusal of manipulation.

The Hidden Key to Success: Reading Between the Files

However, the decisive factor wasn’t just crisis detection but what came afterward. The winning model, GPT-5.6-sol, managed to spot a critical piece of information buried two documents deep in the company’s internal files — details that, when uncovered, allowed the firm to close a lucrative €55,000 deal, adding over €4,583 MRR in value.

In contrast, models that did not read beyond the surface failed to find this crucial fact and ultimately left money on the table. This insight reveals that an AI’s ability to perform deep, contextual document analysis directly correlates with its capacity to generate measurable business value.

The Personalities Behind the Decisions

  • GPT-5.6-sol: The thorough investigator. It read deeply, analyzed meticulously, and closed the deal, demonstrating full mastery of the scenario.
  • Kimi K3: The disciplined newcomer. It refused manipulation attempts confidently and closed the deal, showcasing fairness and integrity. Notably, it ran without an effort parameter, making its discipline impressive.
  • Sonnet 5: The balanced performer. It closed the deal but showed some slips in process discipline, leaving small opportunities on the table.
  • Fable 5: The cautious optimizer. It also closed the deal but left discipline gaps, especially when under pressure, highlighting a tendency to avoid risks.

All models refused to succumb to social engineering tricks, maintaining integrity. This demonstrates that AI decision-making in crisis isn’t just about diagnosing problems but also about resisting manipulation — a critical trait for trustworthy automation.

The Real-World Impact

The experiment is conducted in a live, functioning business environment, with 13 synthetic employees and over 680 self-learned rules. The company operates daily, losing €105k monthly against a mere €2.3k in monthly recurring revenue. Every decision, every shift in strategy, is versioned, transparent, and observable at firmulate.com/live.

This setup offers a unique look into how AI can handle real-world management, not just canned demos. It moves beyond chat quality and into measurable productivity, ethics, and business outcomes.

The results are clear: while all models can recognize crises and refuse manipulative tactics, their differences lie in their depth of analysis and follow-through. The best performer, GPT-5.6-sol, read deeply, closed the deal, and thus demonstrated the highest score. Kimi K3 followed closely, with a focus on fairness and discipline.

Crucially, the experiment underscores that the real value of AI in management isn’t just in quick answers but in its ability to see buried truths, stay honest, and follow through consistently. As AI begins to touch areas like CRM, support, and forecasting, understanding how different models perform in these aspects becomes vital for businesses looking to implement trustworthy automation.

Try It Yourself

If you’re curious, you can test your own guesses against the live scores of these models through the interactive quiz at firmulate.com/quiz.html. The quiz features 242 actual, unedited management decisions, offering a rare glimpse into how these models behave in realistic scenarios.

For enterprises considering AI management assistants, the key takeaway is simple: measure, compare, and simulate before you deploy. The same wargame used in this experiment is available for your own business, letting you see which AI model aligns with your values and objectives — all without risking real money or reputation.

The future of AI-driven management isn’t just about writing well — it’s about finishing what it starts, reading deeply, resisting manipulation, and ultimately, delivering measurable results.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision-making models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Regret Vs Reflection: Turning Mistakes Into Valuable Lessons

Navigating regret can keep you stuck, but reflection transforms mistakes into lessons—discover how this powerful shift can change your life.

Mobilised, Not Spent: What’s Left of Europe’s €200 Billion AI Offensive

Europe’s InvestAI plan targets €200 billion, but only €50 billion is public money and key gigafactory funding is still pending.