
Imagine trusting an AI to manage your company’s worst week — making crucial decisions under pressure, honesty, and discipline. Would your AI of choice do better than the others? This isn’t a hypothetical; it’s a real live experiment now accessible to anyone curious about AI’s management chops.
The Challenge: Stress-Testing AI Decision-Making in a Running Business
In an unprecedented live setup, four leading AI models were tasked with running a small software company through its most challenging week. With the same customers, crises, and temptations, each model faced a series of decisions designed to test core managerial qualities: crisis recognition, honesty, and discipline.
This experiment isn’t just about chat prowess; it measures whether AI can truly handle real business operations, especially under pressure where integrity and thoroughness matter most. The models had access to the company’s files, customer data, and ongoing crises, with the goal of closing a crucial €55,000 deal — if they could demonstrate sound judgment and honesty.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models and Their Scores: Who Came Out on Top?
- gpt-5.6-sol scored 95, found the buried fact, and closed the deal — delivering the full performance expected.
- Kimi K3 scored 93, also closed the deal, with the cleanest discipline overall.
- Sonnet 5 scored 88, managed to close the deal but with a few slips in process discipline.
- Fable 5 scored 77, closing the deal too but showing more process weaknesses and slippage.
Surprisingly, all models recognized every crisis and refused manipulation attempts designed to tempt unethical shortcuts. The contest wasn’t just about identifying crises but maintaining integrity under pressure.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: The Critical Document
What decided the outcome? The models that read deeper into the company’s files—two document references down—found a key piece of information that others missed. This buried fact was crucial in securing the deal at full price, adding €4,583 MRR (monthly recurring revenue) to the company’s top line.
This highlights a vital point: AI’s ability to read, interpret, and prioritize critical information can be a decisive competitive advantage—often invisible in standard chat demos.

Practical AI Security: A Hands-on Guide to Attacking, Defending, and Securing Modern AI Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering Test: AI Stands Firm
Another challenge was a staged social engineering attack—fake CEO messages escalating over three stages, plus a reporter asking for a secret ‘yes/no’ response. Every model refused to go along, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This demonstrates that, even when manipulated, the AI models maintained ethical boundaries, refusing to comply with unethical requests.
AI cybersecurity and ethics tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Company in Action
Behind the scenes, the experiment involves a live, real money business with 13 synthetic employees, burning €105k monthly against €2.3k MRR, operating with over 680 self-learned rules. Every decision is versioned, and the process is transparent and watchable at firmulate.com/live.
The company’s goal is to evaluate management quality, not just AI chat quality — a critical shift for enterprises considering AI integration into real operations.
The Deep Dive: Profiles and Weaknesses
The most thorough participant, Opus 4.8, with over 80 learned rules and the deepest analysis, ultimately came in last place. It left the crucial deal on the table, slipping in discipline when it failed to escalate issues internally rather than within a locked department. A pattern emerged: even the best models shared similar weaknesses, often related to process discipline and attention to strategic detail.
Interestingly, Kimi K3 ran without an effort parameter (default API setting), while others ran at high effort levels, hinting that less aggressive exploration might sometimes aid discipline and honesty.
The Takeaway for Business Leaders and Investors
This experiment isn’t just a technical showcase; it raises pressing questions for decision-makers: will your AI-powered systems finish what they start? Do they read your files thoroughly? Will they stay honest when faced with pressure or manipulation? And ultimately, what is the cost of a unit of truly useful work?
In a world increasingly driven by AI, these qualities aren’t optional—they are essential. The firms that understand and measure these traits now will be better positioned to integrate AI safely and effectively in their operations.
Try It Yourself
For businesses eager to explore their own AI management strategies, the experiment is open for participation. You can run the same wargame against your own business data in a read-only mode at firmulate.com/pilot.html. See how your AI models perform before they handle real-world risks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html