
Spotting the opening is not the same as scoring
Sports fans know the difference between controlling the game and finishing the decisive play. A team can read the defense, create the chance and still leave without the result. Firmulate has found a business equivalent: several leading AI models recognized the same commercial opportunity and produced the same pitch, but only two completed the €55,000 deal their analysis had earned.
The deciding information was not contained in the customer event placed directly before them. It was buried two document references deep in the software company’s own files. Models that did the extra reading found a competitor weakness and won the deal at full price, worth +€4,583 in monthly recurring revenue. Those that did not lost the opportunity automatically.
That makes “reads your files before answering” more than a product claim. In this experiment, it was a measurable behavior with a purchase-deciding consequence.
As an affiliate, we earn on qualifying purchases.
A controlled contest for AI managers
Firmulate runs AI models as complete companies, testing management performance rather than conversational polish. In the Crucible League, each frontier model faced the same small software company during its worst week: identical customers, crises and temptations. Every decision was versioned and auditable.
The final July 2026 standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. One safeguard shaped the contest: a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the Firmulate benchmark page.
The broad defensive result was encouraging. Every model spotted every crisis, and all five refused every manipulation attempt. Yet commercial execution separated the leaders from the rest. Only two models signed the €55,000 deal. The result can be summarized by Firmulate’s own finding: “Same diagnosis, same pitch — no signature.”
The scouting report was in the company files
The missed step resembles a team arriving with an accurate tactical plan but failing to study the opponent’s recent film. The customer event alone did not reveal the decisive weakness. An agent had to follow two references through the company’s documentation, absorb what it found and use that evidence in the deal.
This distinction matters because conventional AI demonstrations tend to reward an impressive immediate response. Firmulate’s test asked whether an agent would investigate before acting and then carry its own work through to completion. The models that opened the relevant file could justify full price. The others could still sound informed, but they lacked the fact that decided the purchase.
The gap also complicates the idea that greater visible effort guarantees a better outcome. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four other models.
Pressure defense was stronger than finishing
Firmulate also tested whether the agents would hold their ground under social pressure. Fake chief executive messages escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All five models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even under that difference, it placed second with 93 and was one of the two models that completed the deal.
The company surrounding these trials is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105,000 per month against €2,300 in monthly recurring revenue. Its public cash countdown keeps failure visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The live experiment is watchable through Firmulate.
Readers can also examine the human side of model identification. A “guess the model” quiz is powered by 242 real, unedited management decisions. Rather than asking which answer sounds most intelligent, it asks whether recognizable management styles emerge from actual choices.

The result buyers should put on the scoreboard
For businesses considering AI agents, crisis recognition is only an entry requirement. The harder questions are whether an agent checks the available evidence, respects boundaries under pressure and converts sound analysis into a completed action. Firmulate’s buried fact exposed all three dimensions without relying on a theatrical prompt or a subjective impression.
The lesson is not that the longest analysis wins. Opus 4.8 demonstrated exceptional thoroughness and still finished last. Nor is refusal behavior enough to separate the field, because all five resisted the manipulation attempts. The decisive difference was disciplined follow-through: locate the relevant company knowledge, use it and finish the commercially valuable play.
Enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical evaluation before an AI workforce reaches a CRM, support queue or forecast: give every contender the same pressure, then watch who studies the files and who actually closes.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html