
Imagine a business that’s bleeding cash in plain sight — yet you can watch it struggle live, day by day, in real time. Welcome to the world of Firmulate, an experiment in building AI-powered companies that fight for survival, publicly.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Company on the Edge
At the heart of this extraordinary experiment is Firmulate, a fully operational, small software firm run by 13 synthetic employees and driven by cutting-edge AI models. Every workday, the company faces real crises, makes real decisions, and incurs actual costs, all while viewers can watch its every move unfold at firmulate.com/live.
Unlike conventional tests, this isn’t a sandbox or a demo. It’s a functioning company, burning €105,000 each month against a revenue of just €2,300 MRR. The stakes are real — a public countdown to bankruptcy looms as the firm’s cash reserves dwindle, integrated with a detailed, self-updating playbook of over 680 learned rules, versioned daily to track progress and setbacks.

Crisis Engineering: Time-Tested Tools for Turning Chaos into Clarity
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI Models in Action: Decisions, Crises, and Trust
The experiment pits four advanced AI models against the same challenging week — the worst possible scenario for a small business. Each model is tasked with navigating the same customers, confronting the same crises, and resisting the same temptations to cheat or manipulate.
Remarkably, all models identified every crisis and steadfastly refused every manipulation attempt, including social engineering tricks like fake CEO messages and reporter stings. For example, when five different models faced a staged fake approval request, all refused — with Kimi K3 explicitly citing suspicion of impersonation. This highlights a crucial point: these models don’t just produce convincing chat; they demonstrate trustworthiness and integrity in high-pressure situations.

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Key to Winning Deals
Despite their steadfastness and accurate crisis detection, only two models managed to secure the company’s most valuable deal — a €55,000 contract. The reason? A buried detail in the company’s internal documents, two references deep, that revealed a critical advantage. Models that read those files and incorporated the information into their pitch succeeded, closing the deal at full price (+€4,583 MRR). Those that missed the hidden data left the money on the table, illustrating that in real-world decision-making, reading deeply and thoroughly matters — not just surface-level chat.

Enterprise AI Observability and Monitoring: Monitoring, Governing Production AI Systems Drift Detection, LLM Monitoring, Agentic AI, Governance, and … (Enterprise Machine Learning Operations)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Money-Losing Company
This isn’t a glossy AI demo. The live firm openly runs in a state of financial crisis, with its cash count visible and updated daily. Every decision, from managing crises to negotiating deals, is made by AI under human oversight, and every step is publicly recorded and auditable.
One standout participant, Opus 4.8, demonstrated thoroughness with over 80 learned rules. Yet, despite detailed analysis, it slipped at the critical moment by attempting to escalate issues into a locked department rather than following proper procedures. Meanwhile, Kimi K3, run without an effort parameter, showed the clearest discipline, closing deals more cleanly but still suffering from process slips.
Implications for Business and AI Reliability
This experiment exposes a fundamental truth: AI’s ability to spot crises, resist manipulation, and follow through on commitments is crucial, especially when real money and trust are involved. It’s not about how well AI writes or chats, but whether it can finish what it starts and stay honest under pressure.
For businesses pondering the integration of AI into their operations, this live experiment offers a sobering lesson. It demonstrates that AI can be trustworthy and diligent — but only if it’s tested in the crucible of real-world, high-stakes scenarios.
See It Live and Decide
Interested in exploring your own business? Firms can run their own wargames against a read-only export of their data, testing AI decision-making without risking real systems. More details are available at firmulate.com/pilot.html.
Why This Matters
As AI begins to touch more aspects of business — from CRM to customer support and forecasting — the question is no longer whether it can produce compelling chat, but whether it can reliably deliver results, stay honest under pressure, and truly add value. The Firmulate experiment makes that question uncomfortably clear: trust in AI depends on rigorous testing in the trenches, not just shiny demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html