
In sport, spotting the opening is only half the job. A team still has to make the pass, take the shot and hold its shape under pressure. Firmulate puts AI models through a business version of that test: the same small software company, the same worst week and the same temptations, with every decision open to replay.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same crisis, different finish
The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. In this contest, partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was not diagnosis or the pitch. It was execution: “Same diagnosis, same pitch — no signature.”
The detail hiding two references deep
The deciding competitor weakness was buried two document references into the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode is a reminder that an opportunity can sit in the material a team already has, while the pressure of a live situation pulls attention elsewhere.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still has to reach the finish
Opus 4.8 produced the most thorough participation, with +80 learned rules and the deepest analyses, yet placed last. It left the deal unsigned and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league offers a useful view of what these models did in this experiment, with that difference in mind.
From watching to trying it on your business
The live company gives the contest a watchable setting. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. The live experiment is available at Firmulate. Its quiz uses 242 real, unedited management decisions and invites readers to guess which model made them.
For a company considering AI agents in its CRM, support queue or forecasts, the next step can be a pilot using a read-only export of its own business. Crisis scenarios run against that company’s context, and the resulting board report can show model rankings and weak points in existing playbooks. Nothing writes back to real systems. That takes the experiment from watching a synthetic company to examining how models handle your own.

To explore a Firmulate pilot for your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
