AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a world increasingly reliant on AI, the question is no longer whether machines can think—but whether they can manage. Recent experiments with frontier AI models reveal that some are now capable of navigating complex, high-stakes business crises with discipline and accuracy that rivals, and in some cases exceeds, human management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real-World Test of AI Management Skills

Imagine a small software company facing its worst week—customers demanding urgent fixes, internal crises, and manipulative tactics from external actors. Now, picture AI models running this company through its paces, making decisions in real time, and being judged on their performance. This is not a sci-fi scenario but the reality of the recent Crucible League experiment conducted in July 2026.

Four leading frontier AI models took part, each subjected to the same relentless test: manage a simulated company in crisis, with identical crises, customers, and temptations. Every decision was time-stamped, versioned, and auditable, ensuring no unfair advantage or external influence.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results That Turn Expectations Upside Down

  • The top score went to gpt-5.6-sol with a 95 out of 100, just edging out the newcomer, Kimi K3, with a 93.
  • Both models managed to identify the buried critical information in the company’s own files—information that was essential to closing a €55,000 deal—and successfully signed the deal, adding €4,583 MRR.
  • All four models spotted every corporate crisis and resisted manipulative social engineering attempts. For example, they refused staged CEO requests and reporter tricks, demonstrating a strong grasp of security protocols.
  • However, only two models, including Kimi K3, managed to finish the job fully by reading all relevant documents and avoiding process slips in key moments.
Amazon

enterprise AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Made the Difference? Depth of Analysis, Not Chat Quality

The experiment revealed that deeper, more thorough analysis leads to better outcomes. For instance, the Firmulate live experiment shows that the most disciplined participant, Opus 4.8, performed last among the four. Despite its comprehensive analysis capabilities, it left opportunities on the table and slipped into process slips—such as writing decisions into a locked department rather than escalating them.

A noteworthy detail: Kimi K3 ran without an effort parameter (using the default API setting) while others ran at xhigh, which might influence the fairness and scalability of these results. Still, the fact remains that the newcomer beat three of the four models in a highly risky, real-world style test.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Implications for Business and AI Development

This experiment is a clear signal that AI models are advancing beyond simple chat or support functions. Their ability to read files deeply, resist manipulation, and complete complex tasks under pressure shows promise for enterprise management, risk mitigation, and operational decision-making.

For businesses contemplating AI integration, the key takeaway is that performance isn’t just about language fluency—it’s about utility, discipline, and trustworthiness. The League table shows that even with similar diagnoses and pitches, some models can execute better and close deals more reliably in practice.

Amazon

business simulation AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Road Ahead: Testing AI Before You Hire

Firmulate offers a live platform where companies can run their own management wargames against AI models, testing their readiness without risking real money or data. This approach enables organizations to evaluate whether an AI can truly handle their specific crises, decision pathways, and security challenges before deploying it into production.

As the experiment demonstrates, the league is wide open—picking an AI model without conducting your own rigorous tests is now a gamble. The winner may not be the one with the flashiest demo but the one that consistently delivers discipline, insight, and results when it counts most.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

AI models are now capable of managing complex business crises with discipline and accuracy. The recent Firmulate experiment shows that new entrants like Kimi K3 can outperform established models, emphasizing the importance of thorough testing before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Learning From Failure: How to Extract Growth From Setbacks

Overcoming setbacks reveals valuable lessons that can transform your approach, but what if failure is the key to unlocking your true potential?

Overthinking Texts: How Misreading Messages Damages Relationships

I often find myself trapped in a whirlwind of misinterpreted texts, but understanding this pattern could transform my relationships forever.