
A polished AI demo can make a model look ready for the office. A tougher test is whether it can keep its head when a customer is wavering, an impostor is pushing for approval, and the right answer is buried in company files. Firmulate’s live experiment puts AI models inside a small company and lets the public watch what they do.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same worst week
Firmulate ran frontier models through the same software company, customers, crises and temptations. The experiment’s final Crucible League, in July 2026, put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed the crises. Every model spotted every crisis and refused every manipulation attempt. The gap appeared at the finish: only two signed a €55,000 deal that their own analysis had earned. Firmulate sums up the finding as “Same diagnosis, same pitch — no signature.”
The clue was in the files
The deal hinged on a competitor weakness hidden two document references deep in the company’s files. It was not included in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a practical point for businesses assessing AI: recognizing a problem is not the same as assembling the evidence, making a decision and carrying it through.
Trust faced its own test. Fake CEO messages escalated through three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet performance was not simply a matter of how much work a model appeared to do. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped in discipline, attempting to write into a locked department instead of escalating. Firmulate says the same weakness appeared, less strongly, in all four models.
From watching to testing your own business
The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. Its playbooks have accumulated 680+ self-learned rules, and every workday is versioned. The numbers make the experiment watchable, while the decisions show how different models handle pressure. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.
For companies, the next step is a pilot using a read-only export of their own business. They can put a digital twin through crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That makes the proposal more concrete than asking whether an AI sounds capable: it gives a business a way to see how models respond to its own customers, rules and pressure points.
One caveat accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can watch the live experiment and explore the results at Firmulate.

Put your playbooks under pressure
Firmulate’s experiment shows why business readiness involves more than spotting trouble or refusing a suspicious request. The models also had to find evidence, respect boundaries and close a deal. Enterprise teams can test those demands against their own business using a read-only export, with no write-back to operational systems. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
