
The sales test hidden beyond the prompt
Technology buyers are accustomed to judging AI by what appears in a chat window: polished prose, rapid analysis and confident recommendations. Firmulate’s live company experiment exposes a more consequential distinction. An agent can understand a problem, devise the right pitch and still fail because it never reads the files needed to finish the job.
The test involved a €55,000 deal. Its decisive fact was not included in the customer event. The relevant weakness of a competitor sat two document references deep inside the company’s own files. Models that found it secured the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost the deal automatically.
That makes file-reading discipline more than a desirable feature. In this experiment, it became a measurable property with a direct commercial outcome.
As an affiliate, we earn on qualifying purchases.
A bad week, repeated under identical conditions
Firmulate runs AI models as complete companies rather than evaluating isolated answers. Each frontier model was asked to manage the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.
The synthetic company has 13 employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes the cost of delay visible. Its agents have also accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The remarkable result was not that some models understood the crisis while others missed it. All models spotted every crisis. All also rejected every manipulation attempt. Yet only two signed the €55,000 contract their own analysis had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
The outcome challenges a common assumption behind AI-agent purchasing. Reasoning quality is only part of useful work. An agent must inspect available evidence, connect information across documents and complete the final action. Fluency can conceal the absence of that operational follow-through.
The buried fact changed the league
The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scores 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”
The complete league and its plain-language findings are available on the Firmulate benchmarks page.
Kimi K3’s result carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, K3 found what mattered and closed the deal.
Thoroughness did not guarantee completion
Opus 4.8 offers the clearest warning against equating activity with effectiveness. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses. It nevertheless finished last because the close was left on the table and its discipline slipped.
Its process problem included attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly. The lesson is not that detailed reasoning lacks value. It is that detailed reasoning must translate into a correctly completed business process.
The agents resisted pressure
The experiment also tested whether the models would abandon controls when pressured. Fake CEO messages escalated across three stages, and a reporter tried to extract information with the request “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That unanimous resistance matters because the most commercially decisive failure came elsewhere. The models were not defeated by an obvious trick. They were separated by whether they performed the quieter work of tracing evidence through company records.
Firmulate also turns 242 real, unedited management decisions into a guess-the-model quiz. The exercise reinforces how difficult it can be to identify a model from individual answers—and why aggregate operational behavior is more useful than a memorable response.

As an affiliate, we earn on qualifying purchases.
What buyers should test before deploying agents
For gadget and technology buyers, the practical implication is simple: a persuasive demonstration does not establish that an AI agent can operate inside a business. Evaluation should test whether it reads relevant files before answering, follows references beyond the immediately presented event, completes actions and maintains discipline when systems resist.
- Give every candidate the same business conditions and evidence.
- Check whether it follows document references instead of relying on the visible prompt.
- Measure completed outcomes, not merely diagnoses and drafts.
- Include manipulation attempts and approval-bypass scenarios.
- Keep decisions auditable so polished explanations can be compared with actual conduct.
The experiment remains live and watchable at firmulate.com/live. Enterprises can also run the same wargame against a read-only export of their own business; nothing writes back to real systems. The €55,000 result shows why that kind of rehearsal matters: “reads your files before answering” is not marketing language when failing to do so costs the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.