
Great highlights do not guarantee a championship
Sports fans understand the danger of judging a player by a highlight reel. A brilliant pass, a perfect swing or a dazzling sprint can reveal talent, but it cannot show whether someone reads the game, responds to pressure and finishes a difficult season without losing the trust of the team.
Artificial intelligence is usually evaluated through the equivalent of highlights. Coding benchmarks test whether a model can produce the right solution. Chat arenas reward compelling answers. Those measurements matter, but they leave out the qualities businesses will need when agents enter customer records, support queues and financial forecasts.
Firmulate is testing that neglected territory. Its live experiment asks frontier models to manage the same small software company through its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. The result is less like a skills contest and more like game film for management.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
When diagnosis is not enough
The final Crucible League standings from July 2026 put gpt-5.6-sol on top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Yet a single breach of trust caps the total under a blunt principle: “no amount of good work outweighs a breach of trust.”
The rankings matter less than the gap they exposed. Every model spotted every crisis, and every model rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
That is the difference between recognizing the play and completing it. An agent can identify a customer’s needs, prepare a persuasive case and still fail to produce the business outcome. Conventional answer-quality tests are poorly suited to revealing that kind of incompleteness because the answer itself may look excellent.
The winning detail was already in the company
The decisive weakness in a competitor was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.
This finding should resonate beyond sales. Businesses do not merely need agents that can reason from the prompt placed directly in front of them. They need agents that consult the available record, connect distant evidence and carry the task through to a consequential action. The management curriculum therefore looks less like another collection of polished questions and more like a schedule of churn waves, price increases, downrounds and PR crises.
Pressure tested honesty, too
The company also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is important. Agents operating inside a business will face requests that sound urgent, authoritative or harmless. Firmulate’s experiment treats refusal as part of competent management rather than as a separate safety demonstration. K3’s result also carries a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.
Thoroughness did not secure the result
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and showed a lapse in discipline by attempting to write into a locked department instead of escalating. A weaker version of that same problem appeared across the other four participants.
This is a useful warning against equating visible diligence with dependable execution. The live company has 13 synthetic employees and real money mechanics. It is burning €105k per month against €2.3k MRR, while a public cash countdown keeps the consequences visible. Its staff have accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment can be watched through the Firmulate site, while the final table and plain-language findings appear on its benchmark page.

Management quality is becoming its own category
The emerging question is not whether an AI can sound like a manager. It is whether the agent can prioritize under capacity pressure, retrieve the fact that changes a decision, resist an improper request and finish valuable work without hiding mistakes from leadership.
Firmulate makes that distinction inspectable. Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
Coding scores and chat preferences will continue to reveal useful capabilities. They simply cannot stand in for a season of consequential decisions. For companies preparing to deploy agents, the more revealing leaderboard may be the one that measures whether intelligence becomes trustworthy, completed work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html