
What a Price War Taught Us About AI Leadership — and Why It Matters for Your Business
Much like running a successful garage or managing an auto shop, overseeing AI systems isn’t just about getting the right words or responses. It’s about how well these systems handle real pressure — especially when the stakes are high, and trust is on the line. Imagine an AI that can’t spot a hidden file or refuses to escalate a problem properly — that’s a gap you won’t see in a chat demo, but it’s critical in the wild.
As an affiliate, we earn on qualifying purchases.
The Real Test: Management, Not Just Conversation
At Firmulate, we ran a groundbreaking live experiment to see how leading AI models perform when managing a real, money-earning software company during its worst week. This wasn’t a simple chat challenge — models faced a series of crises, customer issues, and manipulations, just like a business under fire. Every decision was recorded, auditable, and designed to simulate the pressures of real management.
Scores and Findings
- The top performer, gpt-5.6-sol, scored an impressive 95 out of 100, recognizing every crisis and refusing manipulation attempts. It even identified crucial buried facts in documents, leading to a €55,000 deal.
- Kimi K3, a newcomer, scored just slightly below at 93, and was praised for its discipline in avoiding manipulations. Sonnet 5 and Fable 5 trailed at 88 and 77, respectively, often slipping in process discipline or missing buried info.
- Most models, including Opus 4.8, spotted all the crises and refused manipulations — but only two closed the deal and signed at full price. The rest left money on the table, revealing a gap in management quality, not just chat accuracy.
The Hidden Weaknesses
Despite perfect crisis detection, the decisive difference lay in reading deeper into company files — not just responding to customer events. Models that looked into internal references closed more deals at premium prices, showing that understanding context and trustworthiness is vital.
Manipulation and Trust
Models faced sophisticated social engineering: staged CEO messages and a reporter trick. All five models refused to be fooled — Kimi K3 explained, “Treat the request as a suspected approval-bypass or impersonation.” This demonstrates that AI’s integrity during pressure and deception matters more than just surface-level correctness.
The Live Business Environment
The experiment was run on a real, money-losing company with 13 synthetic employees managing €105k monthly burn against €2.3k MRR. The company’s daily operations, rules, and crises were all live, versioned, and transparent — available at firmulate.com/live. This setup allowed us to observe how AI systems manage ongoing, complex scenarios in real time, not just isolated chat prompts.
Implications for Business
This experiment highlights a critical insight for any company integrating AI: success isn’t just about how well it chatters or responds. Instead, it’s about whether it can finish what it starts, read and trust internal documents, stay honest under pressure, and ultimately deliver measurable, valuable outcomes.
enterprise AI document reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway for Your Garage or Business
If you’re thinking about deploying AI into your operations, whether for customer support, diagnostics, or management, ask yourself: does it do more than just talk? Does it grasp the full context? Will it act honestly when faced with pressure or deception? These qualities are invisible in most demos but are what separate a reliable AI from a risky one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trustworthiness and integrity solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.