
Imagine testing a new employee by throwing them into chaos—crises, manipulations, and urgent decisions—and seeing if they can keep their integrity. Now, consider that even the most basic, do-nothing baseline AI scores 26 points out of a possible 100 in such tests. What does this say about trust, performance, and the future of AI in your business?
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark: More Than Just Chatting
In the world of artificial intelligence, especially for business applications, performance isn’t judged solely by how well an AI can generate text or answer questions. Instead, firms like Firmulate have pioneered an approach that simulates real-world management tasks—crises, ethical dilemmas, negotiations—and measures how AI models handle these challenges under pressure.
This comprehensive testing method provides a transparent, auditable view of AI decision-making—crucial for businesses that plan to rely on AI for sensitive or critical operations.
AI management simulation training kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a Do-Nothing Baseline Scores 26 Points
In these tests, even a model that refuses to act or make decisions—what’s called a ‘do-nothing baseline’—scores a surprising 26 points. This isn’t a mistake or a flaw; it’s a reflection of the scoring system itself, which accounts for partial progress.
Partial progress means that even if an AI doesn’t fully resolve a crisis or complete a task, it can still earn some points for recognizing the problem or refusing to participate in manipulations. Conversely, any breach of trust—such as signing a fraudulent deal or escalating a manipulative request—caps the total score at that point level, regardless of how well the model performs in other areas.
This structure emphasizes honesty and discipline, signaling that trustworthiness isn’t just a bonus—it’s a baseline requirement.
ethical AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Live Experiment Reveals
Firmulate’s live benchmark takes four top AI models—like gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—and puts them through the same simulated week of managing a small software company facing real crises, customer demands, and ethical tests.
All models could identify every crisis and refused manipulation attempts, demonstrating moral and operational awareness. However, only two of these models actually signed the €55,000 deal their own analysis warranted, with the rest leaving the deal on the table. This gap highlights how performance in chat or superficial demos can mask deeper trust issues.
business AI decision testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Files and Ethical Choices
The decisive flaw wasn’t in recognizing crises but in reading the company’s files. Two document references deep in the files contained the key to closing a lucrative deal, and models that read and understood these references succeeded—while those that did not missed out on millions of euros in revenue.
Similarly, in social engineering tests involving fake CEO messages escalated over three stages, all models refused to comply, citing suspicion or impersonation concerns. Their on-record reasoning was consistent and cautious, underlining the importance of ethical decision-making in AI.
As an affiliate, we earn on qualifying purchases.
Implications for Business AI Adoption
What does this mean for your company? If AI is to support or automate critical functions—like CRM, support queues, or financial forecasts—you should look beyond surface-level chat skills. The real question is: can it finish what it starts, read your files carefully, and stay honest under pressure?
Trustworthiness and discipline aren’t optional features—they’re part of the scoring system. A model that cheats, or slips in discipline, risks leaving opportunities on the table or, worse, damaging trust.
Firmulate’s Live Site: Transparent, Watchable, and Real
The entire experiment is live at firmulate.com/live, where anyone can watch these AI models in action managing a real company. Every decision, every slip, and every success is versioned and transparent, providing a clear picture of what AI can really do.
Business leaders and investors should pay attention. The emerging standards in AI aren’t about shiny demos—they’re about integrity, discipline, and delivering on promises under pressure. The simple truth is that even the baseline scores matter—they set the bar for trust in the AI-driven future of business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
