
Coding skill is not management skill
Technology buyers have become accustomed to judging artificial intelligence through coding leaderboards and chat arenas. Those tests can reveal whether a model produces a strong answer. They say far less about whether an agent can prioritize under capacity pressure, follow through across days, resist a dubious instruction or give the board an honest account when the news is bad.
That distinction matters as AI moves beyond the chat window and into customer records, support queues and financial forecasts. An agent may diagnose a problem eloquently and still fail the business. The emerging question is not simply whether the model is intelligent. It is whether the model can manage.
As an affiliate, we earn on qualifying purchases.
A worst week, repeated under controlled conditions
Firmulate, an AI company emulator, turns that question into a live, watchable experiment. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The league table is interesting, but the more revealing result sits beneath it. Every model identified every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s most concise verdict captures the management gap: “Same diagnosis, same pitch — no signature.”
Reading the company was more valuable than reading the room
The decisive fact in that sales contest was not contained in the customer event. It was buried two document references deep inside the company’s own files: a weakness in the competitor’s position. Models that followed the trail won the contract at full price, worth +€4,583 MRR.
This is the kind of difference conventional evaluations can miss. In a chat, the prompt usually brings the relevant context to the model. In a company, the model must decide that context is missing, search for it and recognize which buried detail changes the commercial outcome. A polished response is not enough if the agent fails to do the homework that gives the response leverage.
Security discipline held up better than execution
The pressure did not come only from customers. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was appropriately direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result complicates the familiar fear that capable agents will automatically become reckless under social pressure. In this experiment, the models were consistently alert to manipulation. Their bigger weakness was more mundane: completing legitimate work while respecting operational boundaries.
Opus 4.8 illustrates the point. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its discipline slipped through write attempts into a locked department instead of escalating. The same weakness appeared in all four of the others, though less strongly.
Thoroughness, in other words, did not guarantee decisive or disciplined action. The result should feel familiar to any manager who has watched a talented employee produce excellent analysis while failing to secure the outcome.
A company with consequences, not a scripted demo
The setting gives those decisions weight. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The company can be watched as it operates, making the experiment less like a staged product demonstration and more like an ongoing management record.
There is also an important qualification in the published comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers should keep that difference in mind when interpreting the close finish at the top.
For anyone who wants to examine the outcomes rather than accept a headline, Firmulate publishes its benchmark findings. It also uses 242 real, unedited management decisions in a “guess the model” quiz, a useful reminder of how difficult it can be to identify systems from managerial behavior alone.

business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The category buyers actually need
The scenarios here—churn wave, price increase, downround and PR crisis—look like the beginnings of a new curriculum for AI agents. They test whether a system can triage, investigate, act, escalate and remain candid when choices have consequences beyond the next reply.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is a more relevant pre-employment test than asking an agent to impress in isolation.
Coding quality and conversational quality will remain useful signals. But once an AI agent is entrusted with a company’s work, the decisive measure becomes management quality: whether it reads before acting, finishes what it starts, protects trust and converts good judgment into an actual result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI crisis management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.