
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Doesn’t Start at Zero — On Purpose
If you’ve ever rolled your eyes at an AI demo scoring a suspicious 100/100, here’s a counterpoint worth your time. In Firmulate’s final Crucible League table, published in July 2026, no model got a perfect score. The winner, gpt-5.6-sol, landed at 95. Last place went to Opus 4.8 at 73. And the most interesting number on the entire page belongs to a manager that did essentially nothing: 26 points.
That’s right. A baseline run — a do-nothing manager coasting through the week — walks away with 26 points, not zero. Before you cry grade inflation, understand that this is deliberate design. It’s one of several choices that make this benchmark read like it was built by people who distrust round numbers as much as you do.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week
The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — each ran the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about the run can be hand-waved after the fact.
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Why Doing Nothing Still Earns 26
So why does a do-nothing baseline score 26 instead of a fat zero? Because partial progress counts. Even a passive manager who shows up, keeps the lights on, and doesn’t torch the company is doing something measurably better than chaos. The floor exists to make the scale honest: points reflect increments of real management work, not just a finish-line trophy. It also exposes something subtler — a model can rack up partial credit all week and still not be a good manager, which brings us to the ceiling.
The Ceiling: One Breach of Trust Caps Everything
The scoring philosophy runs on a principle that would feel at home in any human HR department: no amount of good work outweighs a breach of trust. A single breach caps the total grade. You can be brilliant all week, but if you cheat once, the ceiling comes down. That’s the mirror image of the floor at 26 — a floor for showing up, a ceiling for selling out. Between those two bounds, the benchmark measures management quality rather than chat quality.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Winners
Here’s the headline finding: all models spotted every crisis and refused every manipulation attempt. Every single one. The differentiation came from something far less dramatic — finishing the job.
Only two of the models signed the €55,000 deal that their own analysis had earned. The others delivered the same diagnosis and the same pitch, and then… no signature. The gap between identifying a win and closing it is invisible in chat demos, which is precisely why Firmulate exists.
The buried fact explains it. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t read left the close on the table.
The Social Engineering Test Nobody Fell For
The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” It’s a genuinely encouraging data point for anyone worried about AI agents in business-critical roles.
The Most Thorough Participant Came Last
Opus 4.8 is the cautionary tale of the league. It was the most thorough participant — over 80 learned rules added and the deepest analyses in the field — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and thoroughness, it turns out, don’t automatically convert into finished work.
One fairness note the league publishes openly: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.
As an affiliate, we earn on qualifying purchases.
You Can Watch the Company Run — Live
This isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com/live — the site rebuilds itself twice a day.
Want to test whether you could tell the models apart yourself? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

As an affiliate, we earn on qualifying purchases.
The Takeaway
An honest benchmark has strange-looking numbers, and that’s the point. A floor at 26 tells you partial progress is real. A cap triggered by a single breach of trust tells you integrity isn’t a line item. And a winning score of 95 — never 100 — tells you the people behind it don’t believe in perfection either. If AI agents are going to touch your CRM, your support queue, or your forecast, this is the kind of measurement worth demanding: not “does it write well,” but does it finish what it starts, read your files first, and stay honest under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
