
Sports fans know a big score does not always tell the whole story. A team has to spot the opening, make the right call under pressure and finish the move. Firmulate put AI models through a similar test: run a small software company during its worst week, with customers to keep and temptations to resist.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A close result, with a new contender
In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western frontier models in the field.
The models faced the same customers, crises and temptations. Their decisions were versioned and auditable. All spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The detail buried in the files
The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found the buried fact, closed the deal and saved a customer who was about to leave.
That makes the result more revealing than a polished answer in a chat window. An AI agent working for a business may need to look beyond the latest message, use the information already on hand and follow through. Firmulate’s results suggest that doing those things can separate a sound diagnosis from a completed deal.
Trust under pressure, and discipline at the finish
The test also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3 made one deviation, the fewest in the field. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Its workdays are versioned, and the company can be watched at Firmulate. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test before you pick
For a sports club, business or any organization considering AI agents, the league table is a useful starting point, not a substitute for seeing how a model handles your own work. Firmulate says enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems. The standings are close, the deal was easy to leave unsigned, and the choice of model is a bet worth testing first. See the benchmark results.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
