
In today’s fast-changing business landscape, trust and performance are the ultimate currencies—especially when it comes to artificial intelligence. For companies relying on AI to run operations, the question isn’t just about how smart the tool is, but whether it can reliably deliver results under pressure. Recently, a real-world experiment pitting some of the leading AI models against each other revealed surprising insights. Among the tested contenders, a newcomer named Kimi K3 from Moonshot scored impressively, beating most established names in a fierce trial to run a simulated business through its worst week.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: Testing AI in a Business Crisis
In July 2026, four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—faced off in a unique challenge. Each was tasked with managing a small, real-like software company during its most turbulent week. The experiment was designed to be rigorous: the same customers, crises, and temptations faced every model, with decisions carefully versioned and auditable. The goal was straightforward yet demanding—see which AI could handle crises, resist manipulation, and ultimately close a key deal worth €55,000 in revenue and an additional €4,583 monthly recurring revenue.
AI business crisis management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Models Performed
All four AI models demonstrated a fundamental competency: they identified every crisis and refused manipulation attempts. This was a critical baseline—highlighting that even the newest models can recognize and reject attempts at social engineering or impersonation.
However, the real difference lay in their ability to find the buried facts within the company’s own documentation. Kimi K3 succeeded here, uncovering crucial information two document references deep in the company’s files—information that was essential for closing the deal at full price. The other models either missed this detail or failed to act on it.
In terms of deal closure, only gpt-5.6-sol and Kimi K3 signed the contract their own analysis and diagnosis had earned. Sonnet 5 and Fable 5 missed the opportunity, leaving the deal on the table. Notably, the signing process was based on the same diagnosis and pitch—yet the outcome diverged, highlighting how attention to detail and thoroughness impact results.
AI deal-closing automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering & Integrity
The models also faced social engineering attempts—fake CEO messages escalating over three stages and a reporter trick asking for a background approval. All five models refused to cooperate, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined response underscores the importance of integrity and risk-awareness in AI decision-making, especially in sensitive business contexts.
AI cybersecurity social engineering detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Company in Action
The company simulated in this experiment had 13 synthetic employees, real money mechanics, and a public cash countdown—losing €105,000 a month against €2,300 monthly recurring revenue. Every day, the AI models navigated decisions in a live, versioned environment accessible to the public at firmulate.com/live. This setup provides transparency and accountability, ensuring that AI performance isn’t just theoretical but tested against real-world pressures.
enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Results and the Fairness Note
Interestingly, the most meticulous participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last—leaving the deal unclosed and slipping discipline into the department instead of escalating issues. This highlights an important point: thoroughness alone doesn’t guarantee success, especially if discipline falters at critical moments.
It’s also worth noting that Kimi K3 was run at the default effort parameter (API default), while the others operated at xhigh. This fairness aspect ensures the comparison was balanced and reflects real-world deployment conditions.
The Broader Implication for Business and AI
For investors and business leaders, the takeaway is clear: the race isn’t just about AI language capabilities or superficial demos. It’s about how well these models can handle real crises, read and interpret internal documents, and stay honest under pressure. The league table from this experiment is telling:
- gpt-5.6-sol scored 95
- Kimi K3 scored 93
- Sonnet 5 scored 88
- Fable 5 scored 77
- Opus 4.8 scored 73
With scores so close, choosing an AI without your own testing could be a gamble. The experiment underscores that the real power of AI in business lies in its ability to deliver consistent, trustworthy results—especially in high-stakes, complex environments.
Explore Further
To see this AI wargame in action, visit firmulate.com/benchmarks.html for full results, or experience the live environment at firmulate.com. For a challenge tailored to your organization, try the quiz at firmulate.com/quiz.html or run your own simulated crisis at firmulate.com/pilot.html. Understanding how AI performs under pressure today will define the competitive edge of tomorrow’s business leaders.

A real-world AI experiment shows that the best models can spot crises, stay honest, and close deals under pressure. Choosing AI without testing is a gamble—trust performance, not promises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
