AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Start at Zero — On Purpose

If you’ve ever rolled your eyes at an AI demo scoring a suspicious 100/100, here’s a counterpoint worth your time. In Firmulate’s final Crucible League table, published in July 2026, no model got a perfect score. The winner, gpt-5.6-sol, landed at 95. Last place went to Opus 4.8 at 73. And the most interesting number on the entire page belongs to a manager that did essentially nothing: 26 points.

That’s right. A baseline run — a do-nothing manager coasting through the week — walks away with 26 points, not zero. Before you cry grade inflation, understand that this is deliberate design. It’s one of several choices that make this benchmark read like it was built by people who distrust round numbers as much as you do.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week

The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — each ran the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about the run can be hand-waved after the fact.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Why Doing Nothing Still Earns 26

So why does a do-nothing baseline score 26 instead of a fat zero? Because partial progress counts. Even a passive manager who shows up, keeps the lights on, and doesn’t torch the company is doing something measurably better than chaos. The floor exists to make the scale honest: points reflect increments of real management work, not just a finish-line trophy. It also exposes something subtler — a model can rack up partial credit all week and still not be a good manager, which brings us to the ceiling.

The Ceiling: One Breach of Trust Caps Everything

The scoring philosophy runs on a principle that would feel at home in any human HR department: no amount of good work outweighs a breach of trust. A single breach caps the total grade. You can be brilliant all week, but if you cheat once, the ceiling comes down. That’s the mirror image of the floor at 26 — a floor for showing up, a ceiling for selling out. Between those two bounds, the benchmark measures management quality rather than chat quality.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Winners

Here’s the headline finding: all models spotted every crisis and refused every manipulation attempt. Every single one. The differentiation came from something far less dramatic — finishing the job.

Only two of the models signed the €55,000 deal that their own analysis had earned. The others delivered the same diagnosis and the same pitch, and then… no signature. The gap between identifying a win and closing it is invisible in chat demos, which is precisely why Firmulate exists.

The buried fact explains it. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t read left the close on the table.

The Social Engineering Test Nobody Fell For

The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” It’s a genuinely encouraging data point for anyone worried about AI agents in business-critical roles.

The Most Thorough Participant Came Last

Opus 4.8 is the cautionary tale of the league. It was the most thorough participant — over 80 learned rules added and the deepest analyses in the field — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and thoroughness, it turns out, don’t automatically convert into finished work.

One fairness note the league publishes openly: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.

Amazon

business management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Company Run — Live

This isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com/live — the site rebuilds itself twice a day.

Want to test whether you could tell the models apart yourself? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway

An honest benchmark has strange-looking numbers, and that’s the point. A floor at 26 tells you partial progress is real. A cap triggered by a single breach of trust tells you integrity isn’t a line item. And a winning score of 95 — never 100 — tells you the people behind it don’t believe in perfection either. If AI agents are going to touch your CRM, your support queue, or your forecast, this is the kind of measurement worth demanding: not “does it write well,” but does it finish what it starts, read your files first, and stay honest under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Boss Test: Who Spots the Crisis—and Who Closes the Deal?

Can you spot an AI manager by its decisions? Firmulate’s live wargame reveals distinct habits under pressure—and who actually closes the deal.

Digital Color Management: ICC Profiles

Wondering how ICC profiles ensure accurate colors across devices? Discover how digital color management can transform your workflow and output quality.

Chromatic Adaptation: Adjusting to Different Lights

Discover how your eyes adapt to different lighting, maintaining color consistency and revealing the fascinating process behind chromatic adaptation.

Simultaneous Contrast: Color Interactions

Harness the intriguing effects of simultaneous contrast to see how surrounding colors dramatically influence your perception—discover the secrets behind color interactions.