Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Imagine a scenario where a fake CEO tries to manipulate company employees into handing over sensitive information—yet the AI workforce refuses every time. For many, AI security is still a guessing game. But recent live experiments by Firmulate reveal a more promising story: advanced AI models can resist social engineering attempts, even under pressure.

Testing AI Integrity in Real-World Crises

Firmulate recently conducted a groundbreaking live experiment, pitting five leading AI models against a simulated crisis scenario in a real software company. The test was designed to mimic the company’s worst week—full of tempting manipulations, urgent requests, and high-stakes decisions. Every decision was documented, versioned, and auditable, providing a transparent view of AI behavior under stress.

The Crux of the Challenge

The challenge was straightforward yet complex: Could the AI models recognize and resist social engineering tactics, especially as manipulative requests escalated? The scenario involved staged messages from a fake CEO requesting confidential data, and even an attempt to persuade staff to sign off on a deal without proper approval. The models faced repeated pressure to bend rules or bypass safeguards.

The Results: All Models Stayed Honest

Remarkably, all five models identified every crisis and refused every manipulative attempt. This consistency underscores a key insight: integrity under pressure can be tested before a real incident occurs. The models didn’t just provide superficial responses—they refused outright, demonstrating a robust understanding of ethical boundaries.

Beyond the Surface: The Hidden Weakness

While all models performed well on the surface, a deeper analysis revealed a subtle weakness in some. The decisive factor in closing a real deal was linked not just to surface responses but to the reading of internal documentation. Models that delved into the company’s internal files successfully uncovered critical information buried two document references deep—information that, when found, led to securing a €55,000 deal at full price.

Amazon

AI security software for social engineering resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: A Testbed for AI Security

This experiment took place in a company running with 13 synthetic employees engaged in real-money mechanics—burning €105,000 monthly against a modest €2,300 MRR. The environment is publicly accessible at firmulate.com/live, offering transparency and ongoing observation. Every workday, the models are tested against fresh crises, with their decisions logged and available for review.

The Participant Profiles

  • Opus 4.8: Most thorough, analyzing over 80 learned rules, yet ultimately left the close on the table, slipping into department-level escalations.
  • K3: Demonstrated the clearest discipline, running without an effort parameter and refusing manipulations convincingly. As Kimi K3’s on-record reasoning states: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • Sonnet 5: Closed the deal successfully, with minor slips in process discipline.
  • Fable 5: Similar to Sonnet, with slightly more process slips.
  • Baseline: A do-nothing approach scored 26, highlighting how partial progress still counts in assessing AI capabilities. The key takeaway remains: no amount of good work outweighs a breach of trust.
Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

If AI is to touch your CRM, support systems, or forecasting, the question isn’t just about how well it writes or responds in dialogue. It’s about whether it can finish what it starts, stay honest under pressure, and read deeper into your documentation before acting. The experiment shows that these qualities are measurable and can be tested before any real-world crisis hits.

What the Industry Is Learning

  • All five models refused every manipulation attempt, including staged messages and subtle requests.
  • Reading internal files was crucial—those who dove deeper secured the full deal.
  • Performance varies, but even the less disciplined models demonstrated resilience against deception.
  • Integrity isn’t just a theoretical ideal; it’s an operational trait that can be tested and improved in live settings.
Amazon

AI internal documentation analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Businesses

Building trustworthy AI isn’t about isolated chat tests; it requires live, continuous evaluation under realistic pressures. Firmulate offers enterprises a way to run their own wargames against their AI workforce—using a read-only export to simulate crises without risking real systems. Learn more at firmulate.com/pilot.html.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Navigate Life's Pitfalls: Common Mistakes Revealed

Navigate life's pitfalls by uncovering common mistakes that hinder growth—discover the strategies that could transform your journey today!