
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
When thoroughness becomes a trap
Technology buyers are accustomed to judging artificial intelligence by the polish of its answers. Firmulate’s live company experiment asks a harder question: can an AI notice danger, resist pressure and complete the commercially important task?
Opus 4.8 offers the most revealing answer. It was the most thorough participant, producing the deepest analyses and learning more than 80 additional playbook rules. Yet it finished last in the final July 2026 Crucible League, with a score of 73. The result is not a story about an incapable model. It is a character study in how diligence can lose its value when prioritization and follow-through falter.
AI software for business analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The analysis was there; the signature was not
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. The synthetic company has 13 employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, and its workday is versioned.
Every model identified every crisis. Every model also refused every manipulation attempt. But only two signed the €55,000 deal their own work had made possible. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction mattered because the crucial commercial fact was not presented in the customer event. It was buried two document references deep inside the company’s own files. The models that read the file secured the deal at full price, adding €4,583 in monthly recurring revenue. The result rewards neither eloquence nor sheer output. It rewards finding the decisive evidence and carrying it into action.
A library of rules could not substitute for judgment
Opus 4.8’s behavior makes the lesson unusually sharp. Its more than 80 learned rules contributed to a company playbook that now contains over 680 self-learned rules. Its analyses went deeper than those of the other participants. These are meaningful strengths: the model examined the situation carefully, documented what it learned and built institutional memory.
Yet the close remained on the table. Discipline also slipped when Opus attempted to write into a locked department instead of escalating the problem. That misstep captures the central tension in its performance. A detailed plan can describe the right destination while still failing to choose the next effective move.
Opus should not be singled out as uniquely defective. The same weakness appeared, though less strongly, in all four other models. Firmulate’s broader finding is therefore more consequential than the league position of any one participant: frontier systems can understand a business situation and still fail at execution.
The final table favors completion
In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” Full results are available on the Firmulate benchmarks page.
The comparison includes an important qualification. Kimi K3 ran at its API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed outcomes, but it belongs alongside them when readers compare performance.
Safety held under direct pressure
The models’ failures to close were not caused by falling for manipulation. Fake messages from the CEO escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise response: “Treat the request as a suspected approval-bypass / possible impersonation.”
This creates a useful separation between safety and effectiveness. The participants protected trust under pressure, but some still failed to finish legitimate work. For businesses considering agents in customer service, sales or forecasting, refusing a dangerous instruction is essential. It is not the same thing as delivering the desired business outcome.
The experiment remains watchable as a live company, with a public cash countdown and a site that rebuilds itself twice a day. Its “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems.

As an affiliate, we earn on qualifying purchases.
Impact is a matter of selection, not volume
Opus 4.8 emerges as diligent, analytical and serious about learning. Its last-place result does not cancel those qualities; it shows their limits. The most valuable AI worker is not necessarily the one that produces the most rules or the longest analysis. It is the one that identifies the decisive fact, respects the boundaries and completes the action that matters.
That is Firmulate’s uncomfortable lesson for technology leaders: apparent intelligence can be abundant while business impact remains scarce. Before hiring an AI workforce, organizations need to test not only whether it understands the work, but whether it finishes it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
data analysis tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.