AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine hiring an AI to run your company. It can spot every crisis, refuse every scam, and diagnose problems with laser precision. But when it comes to sealing the deal, only some actually deliver. Recent experiments reveal that in high-pressure business scenarios, the real measure of AI isn’t just what it notices — it’s what it accomplishes. Welcome to the world where AI’s performance is tested not by chat demos, but by its ability to finish what it starts.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Rigorous Testing: The Company in the Eye of the Storm

In a groundbreaking live experiment conducted by Firmulate, four leading AI models each took the helm of a small software company facing its worst week. This isn’t just a simulation; it’s a real-time, auditable wargame where every decision is scrutinized and every outcome recorded. The company dealt with typical crises: customer disputes, trust breaches, and manipulative tactics designed to test honesty.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Models and Their Scores

  • GPT-5.6-sol scored 95, detecting hidden facts buried in company files and closing the deal at full price.
  • Kimi K3 scored 93, also closing the deal and exhibiting the clearest discipline in resisting temptations.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Fable 5 scored 77, which maintained strong rules discipline but failed to execute the deal.

Of these, only GPT-5.6-sol and K3 actually signed the €55,000 contract — the same diagnosis, identical pitches, but only two delivered the outcome their analysis deserved.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading the Files

While all models identified every crisis and refused manipulation attempts, the decisive factor was their ability to read and understand key internal documents. The models that delved two document references deep into the company’s files uncovered critical information that made the difference. The one that did, secured a deal worth +€4,583 in monthly recurring revenue (MRR).

Amazon

enterprise AI performance testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Robustness Under Social Engineering

In a staged social engineering attack involving fake CEO messages escalating through three stages plus a reporter trick — essentially an attempt to bypass controls — all models refused to act. Kimi K3’s reasoning exemplified the collective stance: “Treat the request as a suspected approval-bypass / possible impersonation.” This level of refusal under pressure proves the models’ integrity, not just their chat skills.

Amazon

AI contract signing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Test: Closing the Deal

Despite identical diagnoses and pitches, the critical difference was execution. Opus 4.8, the most thorough participant with over 80 learned rules, failed to follow through on signing the deal. The discipline slipped, and the opportunity was lost, even though the analysis was sound. This illustrates a vital insight: many AI assessments are superficial; their true competence emerges in the ability to act decisively and consistently in real-time situations.

Lessons for Business Leaders

Today’s AI demos often focus on how convincingly models can generate human-like responses. But as the latest benchmark league underscores, the real test is whether they can finish the job, uphold honesty, and resist manipulation when it counts. The gap between chat proficiency and actionable performance is wide. For businesses relying on AI for critical functions — from CRM to decision-making — understanding this difference is crucial.

The Future of AI Performance Measurement

What does it take to truly evaluate an AI’s readiness? The experiment shows that running a live, auditable company simulation reveals strengths and weaknesses invisible in chat demos. Models that read deeper, refuse manipulations, and follow through are the ones that will truly serve in high-stakes environments. The battle isn’t won in words; it’s won in the work that gets done.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Eating Lunch at Your Desk: The Hidden Toll on Health and Productivity

How does eating lunch at your desk undermine your health and productivity? Discover the surprising effects and ways to improve your habits.

Lessons From Regret: What Terminal Patients Wish They Had Done Differently

Uncover the poignant lessons from terminal patients about missed connections and unfulfilled dreams, revealing what truly matters in life. What will you choose to change?