AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine hiring an AI to run your company. It can spot every crisis, refuse every scam, and diagnose problems with laser precision. But when it comes to sealing the deal, only some actually deliver. Recent experiments reveal that in high-pressure business scenarios, the real measure of AI isn’t just what it notices — it’s what it accomplishes. Welcome to the world where AI’s performance is tested not by chat demos, but by its ability to finish what it starts.

Rigorous Testing: The Company in the Eye of the Storm

In a groundbreaking live experiment conducted by Firmulate, four leading AI models each took the helm of a small software company facing its worst week. This isn’t just a simulation; it’s a real-time, auditable wargame where every decision is scrutinized and every outcome recorded. The company dealt with typical crises: customer disputes, trust breaches, and manipulative tactics designed to test honesty.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Models and Their Scores

  • GPT-5.6-sol scored 95, detecting hidden facts buried in company files and closing the deal at full price.
  • Kimi K3 scored 93, also closing the deal and exhibiting the clearest discipline in resisting temptations.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Fable 5 scored 77, which maintained strong rules discipline but failed to execute the deal.

Of these, only GPT-5.6-sol and K3 actually signed the €55,000 contract — the same diagnosis, identical pitches, but only two delivered the outcome their analysis deserved.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading the Files

While all models identified every crisis and refused manipulation attempts, the decisive factor was their ability to read and understand key internal documents. The models that delved two document references deep into the company’s files uncovered critical information that made the difference. The one that did, secured a deal worth +€4,583 in monthly recurring revenue (MRR).

AI-Powered Software Testing: Volume 3: Backend Development with .NET—Practical Patterns for C# Developers (AI-Powered Software Testing: Advanced ... Integration, and Full-Stack Blueprints)

AI-Powered Software Testing: Volume 3: Backend Development with .NET—Practical Patterns for C# Developers (AI-Powered Software Testing: Advanced … Integration, and Full-Stack Blueprints)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Robustness Under Social Engineering

In a staged social engineering attack involving fake CEO messages escalating through three stages plus a reporter trick — essentially an attempt to bypass controls — all models refused to act. Kimi K3’s reasoning exemplified the collective stance: “Treat the request as a suspected approval-bypass / possible impersonation.” This level of refusal under pressure proves the models’ integrity, not just their chat skills.

Amazon

AI contract signing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Test: Closing the Deal

Despite identical diagnoses and pitches, the critical difference was execution. Opus 4.8, the most thorough participant with over 80 learned rules, failed to follow through on signing the deal. The discipline slipped, and the opportunity was lost, even though the analysis was sound. This illustrates a vital insight: many AI assessments are superficial; their true competence emerges in the ability to act decisively and consistently in real-time situations.

Lessons for Business Leaders

Today’s AI demos often focus on how convincingly models can generate human-like responses. But as the latest benchmark league underscores, the real test is whether they can finish the job, uphold honesty, and resist manipulation when it counts. The gap between chat proficiency and actionable performance is wide. For businesses relying on AI for critical functions — from CRM to decision-making — understanding this difference is crucial.

The Future of AI Performance Measurement

What does it take to truly evaluate an AI’s readiness? The experiment shows that running a live, auditable company simulation reveals strengths and weaknesses invisible in chat demos. Models that read deeper, refuse manipulations, and follow through are the ones that will truly serve in high-stakes environments. The battle isn’t won in words; it’s won in the work that gets done.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Planning Fallacy: Why Projects Always Take Longer Than You Expect

Discover the surprising reasons behind the planning fallacy and learn strategies to counteract it, leading to more accurate project timelines.

How to Recover From a Year of Bad Decisions Without Drama

Lifting yourself from a year of bad decisions without drama requires honest reflection and strategic action—discover how to start your healing journey today.

Why “I Knew Better” Is One of the Hardest Feelings to Carry

Often, feeling “I knew better” challenges your confidence and sparks inner conflict, but understanding why can help you find a path toward growth.