
Management instincts are becoming a model feature
Choosing an AI model for business used to sound like a specification-sheet exercise: compare capabilities, inspect a demo and pick the system that appears smartest. Firmulate offers a more revealing test. It gives frontier models the same company, the same difficult week and the same opportunities to succeed, stall or break trust.
The result is an unusually tangible technology story. Models that can identify the same crisis and produce similarly persuasive analysis still behave differently when it is time to read the files, follow through and make a consequential decision. Those differences now power an interactive challenge built from 242 real, unedited management decisions. Readers can guess which model made each decision before seeing the answer and its emerging management profile.
This is less like identifying prose style and more like recognizing a colleague by how they work. Does the manager investigate before acting? Does it remain disciplined when a request becomes suspicious? Does it turn a good diagnosis into a completed deal? Firmulate’s wager is that these habits are measurable—and that they matter more than a polished chat response.
As an affiliate, we earn on qualifying purchases.
The same company, under the same pressure
In the Crucible League, each frontier model ran the same small software company through its worst week. Customers, crises and temptations were held constant, while every decision was versioned and auditable. The simulated organization has 13 synthetic employees and unforgiving real-money mechanics: it burns €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its playbook contains more than 680 rules learned through operation.
The final July 2026 table put gpt-5.6-sol in front with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The gap between seeing and finishing
Every model detected every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
The difference was not hidden in the customer interaction. A decisive weakness in the competing offer sat two document references deep inside the company’s own files. Models that followed that trail could close at full price, adding €4,583 in monthly recurring revenue. The episode turns a mundane workplace behavior—reading the available material before acting—into a competitive distinction.
That matters for companies evaluating agents. A system can recognize an opportunity, explain it clearly and even prepare the argument, yet still fail at the final operational step. In a conversational benchmark, the analysis might look excellent. In a running company, the unsigned contract is the outcome.
Pressure revealed discipline as well as judgment
The models also encountered fake CEO messages that escalated over three stages, plus a reporter attempting to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 captured the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s result deserves one methodological note. It ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference, it finished just behind the league leader and displayed the cleanest discipline in the field.
Opus 4.8 illustrates why sheer thoroughness was not enough. It produced the deepest analyses and added 80 learned rules, more than any other participant, but finished last. It left the close on the table and repeatedly tried writing into a locked department instead of escalating the blockage. The same weakness appeared in milder form across the other four participants: recognizing a process obstacle did not always translate into handling it effectively.
A quiz with consequences behind every answer
The quiz makes these contrasts legible without reducing them to abstract scores. Each prompt presents an authentic decision from the experiment and asks the reader to identify its author. The reveal is entertaining, but it also trains attention on a practical question: what kind of manager does a model become when it has context, responsibility and pressure?
Because the decisions are unedited, the personalities arise from behavior rather than branding. The league leader combined investigation with completion. K3 paired suspicion of approval bypasses with strong operational discipline. Opus demonstrated exceptional depth but showed how extensive thinking can coexist with missed execution. The remaining scores show that the field was competitive, while the unsigned deal shows that a small behavioral difference can still dominate the business result.

As an affiliate, we earn on qualifying purchases.
The next model comparison may look more like a wargame
Firmulate’s live experiment suggests that businesses should evaluate AI workers as workers, not merely as answer generators. Crisis recognition and safe refusals are essential, but so are file-reading habits, escalation discipline and the ability to finish valuable work.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing the model’s decisions to be observed without giving the experiment control over production operations.
For gadget and technology enthusiasts, the quiz is the approachable doorway into that larger shift. The challenge is fun because the models appear to have recognizable management personalities. It is important because those personalities can determine whether a company merely receives a strong analysis—or receives the signed deal that analysis was supposed to produce.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
interactive simulation software for training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.