
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A chatbot demo can impress. Running a company is another test.
Moonshot’s Kimi K3 has taken second place in Firmulate’s company-running experiment, finishing ahead of three of four Western frontier models. Its result makes a practical point for businesses considering AI agents: a model’s reputation is no substitute for seeing how it handles your work.
As an affiliate, we earn on qualifying purchases.
Same company, same difficult week
Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. The experiment tracks management decisions rather than chat quality, and each decision is versioned and auditable. The company is a live, watchable experiment, not a fictional scenario.
In the final July 2026 league, gpt-5.6-sol led with 95 points. Kimi K3 followed at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26: partial progress counted, but one breach of trust capped the total.
As an affiliate, we earn on qualifying purchases.
Finding the detail that changes the deal
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The difference came down to a competitor weakness buried two document references deep in company files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
K3 found that buried fact, closed the deal and saved a churning customer. It also resisted all three baits, with only one deviation—the cleanest discipline in the field. In response to the reporter’s “just one yes/no, on background” trick, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
The most thorough participant, Opus 4.8, finished last. It accumulated +80 learned rules and produced the deepest analyses, but left the deal unsigned and slipped on discipline, attempting writes in a locked department instead of escalating. A weaker version of that process weakness appeared in all four models.
As an affiliate, we earn on qualifying purchases.
Useful work under pressure
That gap between recognizing a problem and completing the work is hard to see in a chat demo. Firmulate’s simulated company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can follow the experiment at Firmulate.
The broader question for companies is not only whether an AI agent writes well. If it may touch a CRM, support queue or forecast, can it finish what it starts, read the relevant files and stay honest under pressure? Firmulate’s 242 real, unedited management decisions also power a “guess the model” quiz at firmulate.com. Enterprises can run the wargame on a read-only export of their own business; nothing writes back to real systems.
Fairness note
K3 ran without an effort parameter (API default) while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test before you choose
K3’s second-place finish puts pressure on assumptions about which model can run a business task best. The top two scores were close, but the experiment also shows why a league table alone is not enough: models can reach the same diagnosis and still differ on whether they close the deal. Picking a model without testing it against your own work is a bet.
See the full results and plain-language findings at Firmulate’s benchmark page.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
