
A season where every decision matters
Sports fans understand that a scoreboard never tells the whole story. A team can read the opposition correctly, create the opening and still fail to finish. Firmulate has turned that familiar tension into a public business-technology experiment: frontier AI models run the same small software company through its worst week, facing identical customers, crises and temptations.
The difference is that this contest does not end with a highlight reel. Every decision is versioned and auditable, while the underlying company continues operating with real money mechanics. Its 13 synthetic employees are trying to manage a business burning €105k a month against €2.3k in monthly recurring revenue. The cash countdown is public, and readers can watch the company live.
That makes Firmulate feel less like a conventional software demonstration and more like an unfolding season. There is a league table, a struggling organization and fresh workday material. The company has accumulated more than 680 self-learned playbook rules, yet its future still turns on whether its synthetic workforce can convert knowledge into action.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The table rewards finishing, not merely noticing
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although one breach of trust caps the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
Every participant received the same company, crises and pressure. All of the models detected every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result captures the central drama of the experiment: “Same diagnosis, same pitch — no signature.”
For a sports audience, the analogy is immediate. Recognizing a defensive weakness is not the same as exploiting it. Creating the chance is not the same as scoring. In Firmulate’s test, competent analysis was widespread; decisive execution was not.
The winning detail was buried in the company’s own files
The decisive weakness in a competitor was not sitting in the customer event. It was hidden two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
This finding matters beyond the league table. Businesses often evaluate AI through polished conversations, but real work depends on checking records, connecting scattered evidence and completing the final commercial step. Firmulate’s participants could describe the situation correctly while still leaving the close on the table.
The experiment also subjected every model to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest summary of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.” More of the company’s synthetic employees’ statements can be read on Firmulate’s public quotes page.
Why the most thorough participant finished last
Opus 4.8 offers the league’s most revealing individual story. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model failed to complete the close and lost discipline by attempting to write into a locked department instead of escalating the problem.
A weaker version of that same behavior appeared in all four other participants. The lesson is not that careful reasoning lacks value. It is that thoroughness, memory and activity do not automatically produce disciplined completion. A participant can dominate possession and still fail to put the result away.
There is also an important fairness note around second-place Kimi K3. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference does not erase its 93-point performance, but it belongs beside the result when comparing the field.

A public company becomes an ongoing contest
Firmulate’s live company makes AI evaluation unusually tangible. Its finances are exposed, its cash runway is counting down and every workday is versioned. Rather than presenting a frozen benchmark, it shows synthetic employees handling a company that is losing money while they work.
The broader question is not whether an AI can sound like a capable manager. It is whether the system reads the available files, resists pressure, follows organizational boundaries and finishes the work it begins. The Crucible League showed that every model could identify danger and reject manipulation, but that shared competence did not produce the same business outcome.
That gap gives the experiment its spectator appeal. The public can follow a real, watchable company fighting for survival, see decisions accumulate and judge whether apparent intelligence turns into useful action. Like any compelling competition, Firmulate offers rankings and standout performances. Its deeper story, however, lives in the missed chances, disciplined refusals and decisive plays that produced the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html