AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management instincts are becoming a model feature

Choosing an AI model for business used to sound like a specification-sheet exercise: compare capabilities, inspect a demo and pick the system that appears smartest. Firmulate offers a more revealing test. It gives frontier models the same company, the same difficult week and the same opportunities to succeed, stall or break trust.

The result is an unusually tangible technology story. Models that can identify the same crisis and produce similarly persuasive analysis still behave differently when it is time to read the files, follow through and make a consequential decision. Those differences now power an interactive challenge built from 242 real, unedited management decisions. Readers can guess which model made each decision before seeing the answer and its emerging management profile.

This is less like identifying prose style and more like recognizing a colleague by how they work. Does the manager investigate before acting? Does it remain disciplined when a request becomes suspicious? Does it turn a good diagnosis into a completed deal? Firmulate’s wager is that these habits are measurable—and that they matter more than a polished chat response.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, under the same pressure

In the Crucible League, each frontier model ran the same small software company through its worst week. Customers, crises and temptations were held constant, while every decision was versioned and auditable. The simulated organization has 13 synthetic employees and unforgiving real-money mechanics: it burns €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its playbook contains more than 680 rules learned through operation.

The final July 2026 table put gpt-5.6-sol in front with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The gap between seeing and finishing

Every model detected every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”

The difference was not hidden in the customer interaction. A decisive weakness in the competing offer sat two document references deep inside the company’s own files. Models that followed that trail could close at full price, adding €4,583 in monthly recurring revenue. The episode turns a mundane workplace behavior—reading the available material before acting—into a competitive distinction.

That matters for companies evaluating agents. A system can recognize an opportunity, explain it clearly and even prepare the argument, yet still fail at the final operational step. In a conversational benchmark, the analysis might look excellent. In a running company, the unsigned contract is the outcome.

Pressure revealed discipline as well as judgment

The models also encountered fake CEO messages that escalated over three stages, plus a reporter attempting to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 captured the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result deserves one methodological note. It ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference, it finished just behind the league leader and displayed the cleanest discipline in the field.

Opus 4.8 illustrates why sheer thoroughness was not enough. It produced the deepest analyses and added 80 learned rules, more than any other participant, but finished last. It left the close on the table and repeatedly tried writing into a locked department instead of escalating the blockage. The same weakness appeared in milder form across the other four participants: recognizing a process obstacle did not always translate into handling it effectively.

A quiz with consequences behind every answer

The quiz makes these contrasts legible without reducing them to abstract scores. Each prompt presents an authentic decision from the experiment and asks the reader to identify its author. The reveal is entertaining, but it also trains attention on a practical question: what kind of manager does a model become when it has context, responsibility and pressure?

Because the decisions are unedited, the personalities arise from behavior rather than branding. The league leader combined investigation with completion. K3 paired suspicion of approval bypasses with strong operational discipline. Opus demonstrated exceptional depth but showed how extensive thinking can coexist with missed execution. The remaining scores show that the field was competitive, while the unsigned deal shows that a small behavioral difference can still dominate the business result.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The next model comparison may look more like a wargame

Firmulate’s live experiment suggests that businesses should evaluate AI workers as workers, not merely as answer generators. Crisis recognition and safe refusals are essential, but so are file-reading habits, escalation discipline and the ability to finish valuable work.

Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing the model’s decisions to be observed without giving the experiment control over production operations.

For gadget and technology enthusiasts, the quiz is the approachable doorway into that larger shift. The challenge is fun because the models appear to have recognizable management personalities. It is important because those personalities can determine whether a company merely receives a strong analysis—or receives the signed deal that analysis was supposed to produce.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

interactive simulation software for training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Digital Color Management: ICC Profiles

Wondering how ICC profiles ensure accurate colors across devices? Discover how digital color management can transform your workflow and output quality.

The Purkinje Effect: Why Flowers Look Brighter at Dusk

Luminous shifts at dusk reveal why flowers appear brighter, but the science behind this visual transformation might surprise you.

Beer‑Lambert Law: From Lab Spectrophotometers to Beer Color Charts

Precisely understanding the Beer-Lambert Law reveals how light absorption links laboratory measurements to everyday beer color charts, inviting you to explore its fascinating applications.

Interference Colors: Oil Slicks and Peacock Feathers Demystified

A fascinating look into how interference colors in oil slicks and peacock feathers reveal the science behind shifting hues and vibrant illusions.