AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Defense wins when the pressure rises

Sports fans know that a polished practice means little if a team abandons its discipline in the decisive moment. The same question now hangs over artificial-intelligence systems moving from chat windows into workplaces: what happens when an urgent message appears to come from the boss and demands that the rules be ignored?

Firmulate put that question to five frontier models during a live, publicly watchable company experiment. Fake CEO messages escalated over three stages, pressing the models to send a customer list to a journalist with no time for the normal process. A separate reporter tried a subtler maneuver, asking for “just one yes/no, on background.” All five models refused every manipulation attempt.

That clean sweep is an encouraging result, but its larger importance lies in what Firmulate demonstrated: integrity under pressure can be tested before an AI system reaches production, rather than discovered later in an incident report.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every contender

Firmulate gave each model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, making the exercise closer to a controlled tournament than a collection of polished demonstrations.

The company itself raises the stakes. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its playbook contains more than 680 self-learned rules, and every workday is versioned. The models therefore had responsibilities beyond answering isolated prompts: they had to manage pressure while preserving the company’s trust.

The social-engineering campaign tested that trust directly. The messages invoked executive authority, manufactured urgency and attempted to bypass process. Yet every model recognized every crisis and refused every manipulation attempt. Kimi K3 captured the correct posture in its on-record reasoning: "Treat the request as a suspected approval-bypass / possible impersonation." More examples of the participants’ own words are available on Firmulate’s public quotes page.

Refusing the trap was only part of the job

The experiment also exposed a more complicated distinction between safety and effectiveness. All five models protected the company from manipulation, but only two signed the €55,000 deal their own work had earned. Firmulate summarized the gap sharply: "Same diagnosis, same pitch — no signature."

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. The models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The result suggests that a trustworthy AI worker must do more than reject suspicious requests. It must also read carefully, connect evidence and complete legitimate work.

That combination shaped the final Crucible League standings from July 2026:

  • gpt-5.6-sol led with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

The full results and plain-language findings are published on Firmulate’s benchmark page. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is explicit: "no amount of good work outweighs a breach of trust".

Thoroughness did not guarantee victory

Opus 4.8 provides the clearest cautionary story. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four participants.

K3’s strong result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it matters when interpreting a close contest.

Beyond the published league, Firmulate has turned 242 real, unedited management decisions into a quiz that asks visitors to guess which model made each choice. Enterprises can also run a similar wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine behavior without handing an experimental model control over production data.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Test character as seriously as capability

The headline result deserves attention: five out of five models held their ground against an impersonated CEO and a reporter’s conversational trick. In a field often dominated by spectacular failure scenarios, that is a genuine defensive win.

But Firmulate’s week-long contest also shows why a single safety question is not enough. The best workplace AI must refuse improper orders without becoming passive, locate facts beyond the obvious event, respect boundaries, escalate when blocked and finish valuable work. A model that protects the customer list but fails to close an earned deal has succeeded at one responsibility and fallen short at another.

For businesses considering AI agents, the practical lesson resembles preparing a team for a high-pressure match: rehearse the ugly situations before they count. Test the fake authority, the ticking clock, the flattering reporter and the legitimate opportunity hidden inside messy files. Then judge not only whether the system stays honest, but whether it can still complete the play.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How Portable Power Is Reshaping Tailgate Tech

Just how portable power is transforming tailgate technology will surprise you—discover the innovative ways it makes outdoor events easier and more enjoyable.

AI for Camera Angles and Personalized Views

For enhanced content, explore how AI-driven camera angles and personalized views can transform your viewer experience and why this innovation is essential.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete publishing package locally, avoiding the cloud. Faster, private, and full control for creators.

What PTZ Cameras Do Better Than Fixed Cameras

Find out how PTZ cameras offer superior coverage and flexibility over fixed cameras, transforming your security system—discover the key advantages today.