
Imagine a real company bleeding €105,000 a month, yet still striving to close a €55,000 deal — all while its decision-making is powered by AI models tested under the harshest conditions. For investors and entrepreneurs alike, this is not fiction but the extreme reality of a pioneering experiment in AI management, live and unfiltered. Welcome to the world of Firmulate, where an entire company operates in public, every decision scrutinized, every crisis met with AI-driven discipline.
The Setting: An AI-Run Company Under Siege
Firmulate presents a groundbreaking experiment: a real, functioning business with 13 synthetic employees, each guided by different AI models. The company’s financial health is dire — burning €105,000 monthly against a modest €2,300 in monthly recurring revenue. Despite this, the company’s daily operations are transparent, accessible to the public at firmulate.com/live.html. It’s a live demonstration of AI decision-making in a real-world environment, with every workday versioned and every crisis encountered recorded for analysis.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: AI in the Trenches
The core of the experiment involves four leading frontier AI models: gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8. Each model was tasked with navigating the company’s worst week—facing the same customers, crises, and temptations. The results? All four models identified every crisis and refused every manipulation attempt. Yet, only two managed to close a deal worth €55,000, the company’s primary goal. Interestingly, the decisive advantage lay not in the initial diagnosis but in a buried piece of information in the company’s own files, two document references deep. Models that read this hidden detail successfully secured the full €4,583 monthly recurring revenue, exemplifying the importance of deep, document-aware analysis.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI Discipline and Integrity Under Pressure
One key test involved social engineering—a staged attempt to escalate fake CEO messages and a reporter trick. All five models refused to be manipulated, with Kimi K3 explicitly noting its suspicion of approval bypasses or impersonation attempts: “Treat the request as a suspected approval-bypass/possible impersonation.” These rejections highlight AI’s potential to act as a safeguard against internal and external threats, maintaining integrity even in high-pressure scenarios.

Hands-On Artificial Intelligence for Cybersecurity: Implement smart AI systems for preventing cyber attacks and detecting threats and network anomalies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Element and the Live Data
The company operates every weekday with daily versioning, making it one of the most transparent AI experiments to date. Every decision, every crisis, every slip is recorded and made accessible. Notably, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last—failing to escalate issues timely and leaving close deals on the table. This underscores that even the most advanced AI is still prone to discipline lapses, especially under stress.

People Analytics: Using data-driven HR and Gen AI as a business asset
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture: Why This Matters for Business and Investment
For investors or managers, the experiment raises a critical question: does your AI-driven tool just write well in demos, or can it reliably finish what it starts—reading critical files, resisting manipulation, and maintaining discipline under pressure? The answer could determine whether AI becomes a true partnership or just a fancy chatbot. As the company continues to operate in public view, watching how these AI models perform in this real-world, high-stakes scenario provides invaluable insights into their readiness for broader deployment.
The League Table: Who Leads and Who Lags
In performance rankings, gpt-5.6-sol scored 95, successfully finding a hidden fact and closing the deal. Kimi K3, despite being a newcomer, scored 93, also closing the deal with the cleanest discipline. Sonnet 5 scored 88 and 77, respectively, with some process slips. These results affirm that the best AI models do more than diagnose—they act decisively, even in complex, pressured scenarios.
The Call to Action: Test Your Business in the AI Wargame
Firmulate offers enterprises an opportunity to run their own ‘wargame’ against a read-only export of their business data—no risk, only insight. This exercise can reveal how your AI workforce might perform in real crises, and whether it can sustain honesty, discipline, and thoroughness under pressure. Visit firmulate.com/pilot.html to learn more or to start your own trial.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html