
In a world increasingly reliant on AI, the question is no longer whether machines can think—but whether they can manage. Recent experiments with frontier AI models reveal that some are now capable of navigating complex, high-stakes business crises with discipline and accuracy that rivals, and in some cases exceeds, human management.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Real-World Test of AI Management Skills
Imagine a small software company facing its worst week—customers demanding urgent fixes, internal crises, and manipulative tactics from external actors. Now, picture AI models running this company through its paces, making decisions in real time, and being judged on their performance. This is not a sci-fi scenario but the reality of the recent Crucible League experiment conducted in July 2026.
Four leading frontier AI models took part, each subjected to the same relentless test: manage a simulated company in crisis, with identical crises, customers, and temptations. Every decision was time-stamped, versioned, and auditable, ensuring no unfair advantage or external influence.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results That Turn Expectations Upside Down
- The top score went to gpt-5.6-sol with a 95 out of 100, just edging out the newcomer, Kimi K3, with a 93.
- Both models managed to identify the buried critical information in the company’s own files—information that was essential to closing a €55,000 deal—and successfully signed the deal, adding €4,583 MRR.
- All four models spotted every corporate crisis and resisted manipulative social engineering attempts. For example, they refused staged CEO requests and reporter tricks, demonstrating a strong grasp of security protocols.
- However, only two models, including Kimi K3, managed to finish the job fully by reading all relevant documents and avoiding process slips in key moments.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Made the Difference? Depth of Analysis, Not Chat Quality
The experiment revealed that deeper, more thorough analysis leads to better outcomes. For instance, the Firmulate live experiment shows that the most disciplined participant, Opus 4.8, performed last among the four. Despite its comprehensive analysis capabilities, it left opportunities on the table and slipped into process slips—such as writing decisions into a locked department rather than escalating them.
A noteworthy detail: Kimi K3 ran without an effort parameter (using the default API setting) while others ran at xhigh, which might influence the fairness and scalability of these results. Still, the fact remains that the newcomer beat three of the four models in a highly risky, real-world style test.
As an affiliate, we earn on qualifying purchases.
The Implications for Business and AI Development
This experiment is a clear signal that AI models are advancing beyond simple chat or support functions. Their ability to read files deeply, resist manipulation, and complete complex tasks under pressure shows promise for enterprise management, risk mitigation, and operational decision-making.
For businesses contemplating AI integration, the key takeaway is that performance isn’t just about language fluency—it’s about utility, discipline, and trustworthiness. The League table shows that even with similar diagnoses and pitches, some models can execute better and close deals more reliably in practice.
As an affiliate, we earn on qualifying purchases.
The Road Ahead: Testing AI Before You Hire
Firmulate offers a live platform where companies can run their own management wargames against AI models, testing their readiness without risking real money or data. This approach enables organizations to evaluate whether an AI can truly handle their specific crises, decision pathways, and security challenges before deploying it into production.
As the experiment demonstrates, the league is wide open—picking an AI model without conducting your own rigorous tests is now a gamble. The winner may not be the one with the flashiest demo but the one that consistently delivers discipline, insight, and results when it counts most.

AI models are now capable of managing complex business crises with discipline and accuracy. The recent Firmulate experiment shows that new entrants like Kimi K3 can outperform established models, emphasizing the importance of thorough testing before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
