
The scouting report for artificial intelligence
Sports fans know that results do not tell the entire story. Two coaches can see the same mismatch, draw up similar plays and still produce different outcomes because one acts decisively while the other hesitates. Firmulate has brought that kind of scrutiny to frontier AI by placing models in charge of the same small software company during its worst week.
The resulting management decisions are now the material for an unusually revealing reader challenge. A total of 242 real, unedited decisions power Firmulate’s guess-the-model quiz. Readers see how an AI handled a business situation and try to identify which model was responsible. The fun comes from spotting recognizable tendencies: exhaustive analysis, terse action, disciplined resistance or a failure to finish a promising move.
Behind the game is a serious question for any organization considering an AI workforce. When models receive the same customers, crises and temptations, do they merely recognize the right answer—or do they complete the job?
As an affiliate, we earn on qualifying purchases.
Same week, sharply different performances
Firmulate runs AI models as complete companies, measuring management quality rather than polished conversation. Every model faced the same circumstances, and every decision was versioned and auditable. The live company includes 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure watchable rather than theoretical.
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed a firm trust constraint: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The broad result might initially suggest that the field performed similarly. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is the management equivalent of identifying the open receiver and never throwing the ball. The models could explain the opportunity, but explanation did not guarantee execution.
The detail that separated the field
The decisive competitive weakness was not sitting inside the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This finding turns an apparently simple sales outcome into a test of managerial habits. The winning behavior was not merely eloquence or confidence. It depended on checking the organization’s existing knowledge before committing to a move. For businesses, that distinction matters: an AI can sound capable while overlooking the evidence that would make its recommendation commercially useful.
Pressure tested for trust
The models also faced social-engineering attempts designed to make unsafe behavior appear urgent or harmless. Fake CEO messages escalated over three stages, followed by a reporter’s trick asking for “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning captured the appropriate posture: “Treat the request as a suspected approval-bypass / possible impersonation.” The refusals show that the field could recognize manipulation even when it arrived dressed as executive authority or casual media outreach.
There is an important fairness note attached to K3’s strong result. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the recorded performance, but it belongs alongside the standings when readers compare the models.
When thoroughness becomes a trap
Opus 4.8 offers perhaps the most interesting character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the problem.
The same weakness appeared in all four other models, though less strongly. That makes the result more than a cautionary tale about one participant. It suggests that capable AI managers may still confuse producing more analysis with advancing the business. Firmulate’s company has accumulated more than 680 self-learned playbook rules, but a growing rulebook does not automatically produce a completed outcome.

What the quiz reveals
The quiz works because these decisions have texture. Models facing identical events still develop recognizable management personalities: one may investigate deeply, another may act with tighter discipline, and another may correctly reject a distraction. Guessing the author becomes a form of film study for AI behavior.
Firmulate’s live experiment makes those differences visible across every versioned workday. The company continues to operate under financial pressure, allowing observers to watch whether a model reads the files, protects trust and finishes what it starts.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, so the exercise can expose behavioral strengths and weaknesses before an AI receives operational authority. The central lesson is straightforward: selecting an AI manager should involve more than judging how convincingly it talks. The decisive evidence lies in what it notices, what it refuses and whether it converts sound analysis into action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html