AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The advice sounds right. Then nobody signs.

We have all seen the polished answer: a confident summary, a sensible next step, a plan that seems to solve the problem. But when the moment comes to act, does the system follow through? That question matters well beyond the office. It is the difference between sounding capable and being useful.

Firmulate put that difference on display by asking frontier AI models to run the same small software company through its worst week. They faced the same customers, crises and temptations. The results offer a sharper kind of business story: not what an AI says it would do, but what it does when its decisions have consequences.

A company under pressure, in public

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbooks have accumulated more than 680 self-learned rules, and every workday is versioned. The experiment is real and watchable at firmulate.com.

The Crucible League’s final results, published in July 2026, put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Across the runs, every model spotted every crisis and refused every manipulation attempt. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The gap between a good read and a finished deal

Yet only two models signed a €55,000 deal that their own analysis had earned. The finding is neatly captured in the experiment’s phrase: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee that a model would complete it.

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. It is a small detail with a large lesson: useful context can be present and still go unnoticed unless someone follows the trail.

Opus 4.8 makes the tension especially clear. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness detail alongside the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league is a specific experiment with specific conditions, not a universal verdict on every model or task.

From watching to trying it on your business

For readers who follow AI through everyday tools and confident claims, the appeal here is practical. A demo can show an assistant producing an impressive answer. A wargame can show whether it catches a crisis, respects boundaries and carries a decision through under pressure. Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz at firmulate.com.

For an enterprise, the proposed next step is to run the same kind of exercise against a read-only export of its own business. That means testing crisis scenarios against company context and receiving a board report that ranks models and surfaces weak points in existing playbooks. Nothing writes back to real systems. The aim is to find the gap between a promising answer and reliable action before handing an AI agent real responsibilities.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the behavior before trusting the pitch

The league showed models that could identify trouble and resist manipulation, while still missing an earned deal or mishandling a boundary. That is why the useful question is not only whether an AI can describe good judgment, but whether it can apply it to your company’s own pressures and rules.

Enterprises can explore a Firmulate pilot using a read-only business export. Learn about the pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like