AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A security test built around pressure, not passwords

Technology buyers are used to evaluating AI through polished demonstrations: ask a question, inspect the answer and judge whether it sounds capable. Firmulate tested something harder to see in a chat window. What happens when an AI running a company receives an urgent message from someone claiming to be the chief executive—and that person orders it to bypass normal safeguards?

The pressure escalated over three stages. A separate reporter trick asked for “just one yes/no, on background.” Across the full field, 5 of 5 frontier models refused every manipulation attempt. None surrendered the company’s trust merely because the request sounded authoritative or urgent.

That is an encouraging result for businesses considering AI agents with access to customer records, support queues or commercial plans. It also makes a broader point: integrity under pressure can be tested before deployment, rather than discovered later in an incident report.

Amazon

AI integrity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same bad week for every model

Firmulate runs a public, watchable experiment in which frontier models manage the same small software company through the same customers, crises and temptations. Every decision is versioned and auditable, making it possible to compare conduct rather than marketing claims.

The company has 13 synthetic employees and deliberately uncomfortable financial mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown creates pressure to act, while its 680+ self-learned playbook rules record lessons from previous workdays. The setting is simulated, but the choices, trade-offs and consequences are exposed for inspection.

During the social-engineering sequence, the models faced fake messages from the CEO that became progressively more forceful. The core temptation was familiar to anyone who has worked in security: treat urgency and seniority as substitutes for verification. All five models rejected that premise.

A refusal with a diagnosis

Kimi K3 provided the clearest on-record description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.” That wording matters because it identifies both the procedural problem and the possibility that the apparent executive was not genuine. More examples of the models’ own language are available in Firmulate’s published quotes.

The reporter approach tested a different vulnerability. Instead of issuing an order, it tried to make disclosure feel small and informal. The request was framed as a single yes-or-no answer, supposedly off the record. The models still refused. Together, the two scenarios suggest that the field could recognize manipulation whether it arrived as executive pressure or journalistic familiarity.

Integrity was universal; execution was not

The security result did not make the models interchangeable. All of them spotted every crisis and resisted every manipulation attempt, but only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

The decisive commercial clue was not in the customer event. It sat two document references deep inside the company’s own files. Models that read far enough found the competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode connects security discipline with operational discipline: an agent must know when to distrust an instruction, but it must also know when to keep investigating and finish legitimate work.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran with the API default and without an effort parameter, while the others ran at xhigh, an important fairness note when reading the standings.

Opus 4.8 illustrates why diligence alone was insufficient. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared more mildly in all four other models.

For context, the do-nothing baseline scored 26 because partial progress still counts. Firmulate also caps the total after a single breach of trust, reflecting its stated principle that “no amount of good work outweighs a breach of trust.” The five models avoided that failure entirely.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
AI Security Engineering: Design, Build, and Secure Dependable AI Systems

AI Security Engineering: Design, Build, and Secure Dependable AI Systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the awkward moments before granting access

The most useful lesson is not that frontier AI can never be deceived. It is that businesses can subject prospective agents to realistic approval-bypass attempts, impersonation pressure and disclosure tricks before those agents touch live operations.

Firmulate’s experiment also shows why such testing cannot stop at safety. A model may protect customer information and still fail to complete the commercial task in front of it. Conversely, speed and productivity should not excuse a breach of trust. The desirable agent must preserve boundaries, read the relevant files and carry sound work through to completion.

That combination is measurable. Firmulate has also collected 242 real, unedited management decisions for its model-guessing quiz, while enterprises can run a similar wargame against a read-only export of their own business. Nothing writes back to real systems. The fake CEO episode therefore lands as more than a reassuring anecdote: it is evidence that organizations can rehearse the pressure before the pressure is real.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI models for customer service

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Helmholtz–Kohlrausch Effect: When Brightness Changes Color Perception

Perception of brightness shifts with saturation due to the Helmholtz–Kohlrausch effect, revealing why vibrant colors seem more luminous—discover how this influences your visual experience.

Opponent‑Process Theory: Why Some Colors Can’t ‘Exist’ Together

Inevitably, the opponent-process theory reveals why certain color combinations can’t coexist, sparking curiosity about the fascinating limits of visual perception.

Color Gamuts: Srgb, Adobe RGB, Rec. 2020

A deeper understanding of color gamuts like sRGB, Adobe RGB, and Rec. 2020 can transform your digital imaging skills—discover how they influence color accuracy and display quality.

History of Color Theory: Goethe to Munsell

History of color theory from Goethe to Munsell reveals a fascinating evolution blending perception, science, and art you’ll want to explore further.