firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In today’s fast-changing business landscape, trust and performance are the ultimate currencies—especially when it comes to artificial intelligence. For companies relying on AI to run operations, the question isn’t just about how smart the tool is, but whether it can reliably deliver results under pressure. Recently, a real-world experiment pitting some of the leading AI models against each other revealed surprising insights. Among the tested contenders, a newcomer named Kimi K3 from Moonshot scored impressively, beating most established names in a fierce trial to run a simulated business through its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI in a Business Crisis

In July 2026, four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—faced off in a unique challenge. Each was tasked with managing a small, real-like software company during its most turbulent week. The experiment was designed to be rigorous: the same customers, crises, and temptations faced every model, with decisions carefully versioned and auditable. The goal was straightforward yet demanding—see which AI could handle crises, resist manipulation, and ultimately close a key deal worth €55,000 in revenue and an additional €4,583 monthly recurring revenue.

Amazon

AI business crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Performed

All four AI models demonstrated a fundamental competency: they identified every crisis and refused manipulation attempts. This was a critical baseline—highlighting that even the newest models can recognize and reject attempts at social engineering or impersonation.

However, the real difference lay in their ability to find the buried facts within the company’s own documentation. Kimi K3 succeeded here, uncovering crucial information two document references deep in the company’s files—information that was essential for closing the deal at full price. The other models either missed this detail or failed to act on it.

In terms of deal closure, only gpt-5.6-sol and Kimi K3 signed the contract their own analysis and diagnosis had earned. Sonnet 5 and Fable 5 missed the opportunity, leaving the deal on the table. Notably, the signing process was based on the same diagnosis and pitch—yet the outcome diverged, highlighting how attention to detail and thoroughness impact results.

Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Handling Social Engineering & Integrity

The models also faced social engineering attempts—fake CEO messages escalating over three stages and a reporter trick asking for a background approval. All five models refused to cooperate, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined response underscores the importance of integrity and risk-awareness in AI decision-making, especially in sensitive business contexts.

Amazon

AI cybersecurity social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Company in Action

The company simulated in this experiment had 13 synthetic employees, real money mechanics, and a public cash countdown—losing €105,000 a month against €2,300 monthly recurring revenue. Every day, the AI models navigated decisions in a live, versioned environment accessible to the public at firmulate.com/live. This setup provides transparency and accountability, ensuring that AI performance isn’t just theoretical but tested against real-world pressures.

Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Results and the Fairness Note

Interestingly, the most meticulous participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last—leaving the deal unclosed and slipping discipline into the department instead of escalating issues. This highlights an important point: thoroughness alone doesn’t guarantee success, especially if discipline falters at critical moments.

It’s also worth noting that Kimi K3 was run at the default effort parameter (API default), while the others operated at xhigh. This fairness aspect ensures the comparison was balanced and reflects real-world deployment conditions.

The Broader Implication for Business and AI

For investors and business leaders, the takeaway is clear: the race isn’t just about AI language capabilities or superficial demos. It’s about how well these models can handle real crises, read and interpret internal documents, and stay honest under pressure. The league table from this experiment is telling:

  • gpt-5.6-sol scored 95
  • Kimi K3 scored 93
  • Sonnet 5 scored 88
  • Fable 5 scored 77
  • Opus 4.8 scored 73

With scores so close, choosing an AI without your own testing could be a gamble. The experiment underscores that the real power of AI in business lies in its ability to deliver consistent, trustworthy results—especially in high-stakes, complex environments.

Explore Further

To see this AI wargame in action, visit firmulate.com/benchmarks.html for full results, or experience the live environment at firmulate.com. For a challenge tailored to your organization, try the quiz at firmulate.com/quiz.html or run your own simulated crisis at firmulate.com/pilot.html. Understanding how AI performs under pressure today will define the competitive edge of tomorrow’s business leaders.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

A real-world AI experiment shows that the best models can spot crises, stay honest, and close deals under pressure. Choosing AI without testing is a gamble—trust performance, not promises.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Set Approval Threshold Alerts for Payment Operations

Meta description: “Master how to set approval threshold alerts for payment operations to effectively detect anomalies and safeguard your transactions, ensuring your system stays secure and efficient.

2D Barcode Scanners: When You Need More Than Basic UPC Reading

Discover the top 2D barcode scanners for small businesses in 2026. Find reliable, versatile options like Eyoyo EYH2, Tera wireless models, and more.

Document Scanners for Accounting Offices: What Finance Teams Need Most

Discover the top document scanners for accounting offices in 2026. Find the best overall, portable, and budget-friendly options tailored for professionals.

How to Build a Better Escalation Path for Payment Issues

Payments can be complex—discover how to build a better escalation path to ensure quick, transparent resolutions and keep customer trust intact.