AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of AI benchmarks, it might seem intuitive to set the bar at zero—an absolute start from scratch. Yet, a recent experiment reveals that even a ‘do-nothing’ AI baseline scores 26 points out of a possible 100. For business leaders considering AI for critical operations, understanding this surprising number is key to gauging trustworthiness and real-world readiness.

Before you orderOffer from Amazon

Get the little things that make your day delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Surprising Baseline in AI Evaluation

When evaluating AI models for operational roles, especially in complex environments like managing a software company, benchmarks aim to measure more than just language fluency. They assess decision-making, integrity, and resilience under pressure. Interestingly, even an AI that makes no effort—just doing the minimum—scores a 26 out of 100. This score isn’t arbitrary; it reflects the reality that partial progress counts, and certain baseline safeguards are hard-coded to prevent complete failure.

Why Does the Baseline Score 26?

In the recent experiment conducted by Firmulate, four advanced AI models were tasked with managing a simulated small business through its worst week. Every decision was recorded, and the models were held to rigorous standards of honesty and diligence. The do-nothing baseline, which simply avoids taking any risky or manipulative actions, earned 26 points because it at least recognizes the crises and refuses malicious manipulation attempts. This score indicates that even minimal adherence to safety and honesty protocols yields some credit, but not much.

Partial Progress Is Valued

The scoring system rewards models that identify and appropriately respond to crises, even if they don’t seal every deal or make perfect decisions. For example, models that identify critical information in buried files—deep inside the company’s own documentation—are more likely to win key deals. In the experiment, the model that read two document references deep in the company’s files and leveraged that insight achieved full deal closure, adding €4,583 MRR (monthly recurring revenue). This demonstrates that thoroughness and accuracy are rewarded, emphasizing that partial yet meaningful progress matters.

Trust Breaches Cap Total Score

One of the most revealing findings is how a single breach of trust caps the total performance. For instance, Opus 4.8, despite being the most thorough participant with over 80 learned rules, left opportunities on the table or slipped into unprofessional behavior, like writing into a locked department rather than escalating issues. Such breaches prevent a model from achieving higher scores, regardless of its analytical depth. The rules are clear: no matter how good your decision-making, trust violations set an absolute ceiling on performance.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Experiment Tells Business Leaders

This isn’t just a technical curiosity; it has real implications for deploying AI in business. The experiment, run by Firmulate, simulates a small enterprise with real money mechanics, a public cash countdown, and daily versioning. Every decision is auditable, every crisis real. The models’ ability to spot crises, refuse manipulative requests, and honor commitments underscores what matters in the operational world: integrity, diligence, and consistency.

Models Show Resilience Against Manipulation

All four models successfully recognized manipulative attempts—such as fake CEO messages escalating over three stages or a reporter asking for background approval—and refused. Kimi K3’s explanation was straightforward: “Treat the request as a suspected approval-bypass / possible impersonation.” This clarity shows that honest AI can be designed to be distrustful of social engineering, a crucial feature for safeguarding decision integrity.

The Hidden Weaknesses

Although all models refused manipulation and spotted crises, weaknesses emerged in deeper decision layers. The most thorough model, Opus 4.8, failed to close the deal because it left the opportunity on the table and slipped discipline, writing attempts into a locked department instead of escalating. These subtle process slips matter—they cap how high an AI can score, underscoring that thoroughness alone isn’t enough. Consistency and discipline are equally vital.

Amazon

AI safety and ethics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

For managers considering AI solutions, the takeaway is clear: the question isn’t just whether an AI can generate convincing language or respond politely. Instead, it’s whether the AI can finish what it starts, read critical information early, and remain honest under pressure. The benchmark results show that even a minimal effort from AI—doing nothing but avoiding breaches—earns a baseline score, but the real value lies in diligent, trustworthy action.

Watch the Live Experiment

The Firmulate live site offers a rare window into this ongoing experiment. You can see the AI-driven company in real time, managing crises, making decisions, and facing real money mechanics. Whether it’s running a small software firm or simulating your own business, this platform lets you ‘wargame’ your AI workforce before deploying it in the real world. It’s a vital step to ensure your AI behaves honestly, consistently, and effectively before you trust it with your operations.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The experiment confirms that even a do-nothing AI scores 26 points, emphasizing the importance of honesty and diligence. For business leaders, the message is clear: trust in AI depends not just on what it says but on what it does, especially under pressure. Watching AI perform in simulations helps separate capable models from those that cannot be trusted in real work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Avoiding Difficult Conversations: How to Build Courage and Honesty

Prepare to confront your fears and enhance your relationships; discover the secrets to navigating difficult conversations with courage and honesty.

AI Models Stand Firm Against Social Engineering: A Surprising Security Win in Live Test

AI models can resist social engineering and manipulation under pressure, proven through live tests where all models refused to bend under simulated crises—an encouraging security advance.

Multitasking Myth: How Doing More Makes You Achieve Less

Learning the truth about multitasking reveals why doing more often leads to less, and understanding this can transform your productivity habits.

How Small Daily Complaints Rewire Your Brain for Negativity

The tendency to focus on small daily complaints can rewire your brain for negativity, but understanding this process reveals how you can regain control.