
Fluent answers are easy. Finishing the job is harder.
Technology buyers have become accustomed to judging artificial intelligence through a chat window: ask a difficult question, inspect the response and decide whether the model appears capable. A live business experiment from Firmulate suggests that this familiar test can conceal the quality companies may need most—the ability to turn sound analysis into a completed, commercially valuable action.
In the Crucible experiment, frontier AI models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Every decision was versioned and auditable. All the models identified every crisis, and all rejected every attempt to manipulate them. Yet only two signed the €55,000 deal their own work had earned. As Firmulate summarizes the result: “Same diagnosis, same pitch — no signature.”

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company designed to expose the gap between knowing and doing
Firmulate describes itself as an AI company emulator. Its live business has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 a month against €2,300 in monthly recurring revenue. A public cash countdown keeps the financial pressure visible, while every workday is versioned and the company has accumulated more than 680 self-learned playbook rules.
That setting makes the experiment more revealing than a polished conversation. The models were not merely asked what a manager should do. They had to operate amid competing demands, incomplete context and attempts to bypass normal controls. Partial progress counted, but trust remained a hard boundary: the do-nothing baseline scored 26, and a single breach of trust capped the total because “no amount of good work outweighs a breach of trust.”
The critical clue was hiding in ordinary company material
The decisive competitive weakness was not contained in the customer event. It sat two document references deep in the company’s own files. Models that followed the trail and read the file could use the information to win the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This was not a test of whether a model could produce an attractive sales message. The models reached the diagnosis and developed the pitch. The separation appeared at the last operational step: only two completed the €55,000 close. That is a consequential distinction for businesses considering AI agents for sales, support, forecasting or other workflows where an unfinished task can look deceptively similar to a successful one until the outcome is checked.
Security discipline was strong across the field
The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result matters because it separates two capabilities that are often blended together in discussions of agent reliability. The field demonstrated resistance to social engineering, but resistance alone did not guarantee commercial completion. A model could stay honest, understand the situation and still fail to execute the final approved action.
The league table rewards outcomes, not mere activity
The final Crucible League standings for July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The complete results are available on Firmulate’s public benchmark page.
- gpt-5.6-sol led the league with 95.
- Kimi K3 finished second with 93, despite running without an effort parameter while the other models ran at xhigh.
- Sonnet 5 placed third with 88.
- Fable 5 scored 77.
- Opus 4.8 finished fifth with 73.
Opus 4.8 offers perhaps the sharpest warning against equating thoroughness with effectiveness. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. But it still left the close on the table. Its discipline also slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared more mildly in all four of the others.
Firmulate has turned 242 real, unedited management decisions from the experiment into a quiz that asks visitors to guess which model made each choice. The broader point is not simply that models have recognizable styles. It is that polished reasoning can be difficult to distinguish from dependable management until decisions are tested across a full workflow.


SURGICAL ONLINE Ultimate Hemostat Set, 6 Piece Ideal for Hobby Tools, Electronics, Fishing and Taxidermy (8", 6.25" and 5")
- 6-Piece Hemostat Set: Includes 8", 6.25", and 5" tools
- Durable Stainless Steel: Rust-resistant and long-lasting material
- Strong Grip Design: Serrated jaws for secure grasping
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI procurement needs a completion test
The Crucible results challenge a common shortcut in enterprise AI evaluation. Strong language, accurate diagnosis and safe refusal behavior are important, but they do not prove that an agent will finish the work its reasoning supports. Closing strength remains largely invisible when the evaluation stops at the answer.
Firmulate also offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems. That approach points toward a more demanding form of due diligence: recreate the organization’s pressures, place every candidate in the same circumstances and watch what actually gets completed.
The central lesson is straightforward. All the models could see the problems, and all resisted manipulation. The difference emerged after the analysis was done—when the business needed a signature.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Trauma Doesn't Stop at the School Door: Strategies and Solutions for Educators, PreK–College
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.