AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A chatbot demo can impress. Running a company is another test.

Moonshot’s Kimi K3 has taken second place in Firmulate’s company-running experiment, finishing ahead of three of four Western frontier models. Its result makes a practical point for businesses considering AI agents: a model’s reputation is no substitute for seeing how it handles your work.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, same difficult week

Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. The experiment tracks management decisions rather than chat quality, and each decision is versioned and auditable. The company is a live, watchable experiment, not a fictional scenario.

In the final July 2026 league, gpt-5.6-sol led with 95 points. Kimi K3 followed at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26: partial progress counted, but one breach of trust capped the total.

Amazon

business AI chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the detail that changes the deal

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The difference came down to a competitor weakness buried two document references deep in company files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3 found that buried fact, closed the deal and saved a churning customer. It also resisted all three baits, with only one deviation—the cleanest discipline in the field. In response to the reporter’s “just one yes/no, on background” trick, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most thorough participant, Opus 4.8, finished last. It accumulated +80 learned rules and produced the deepest analyses, but left the deal unsigned and slipped on discipline, attempting writes in a locked department instead of escalating. A weaker version of that process weakness appeared in all four models.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Useful work under pressure

That gap between recognizing a problem and completing the work is hard to see in a chat demo. Firmulate’s simulated company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can follow the experiment at Firmulate.

The broader question for companies is not only whether an AI agent writes well. If it may touch a CRM, support queue or forecast, can it finish what it starts, read the relevant files and stay honest under pressure? Firmulate’s 242 real, unedited management decisions also power a “guess the model” quiz at firmulate.com. Enterprises can run the wargame on a read-only export of their own business; nothing writes back to real systems.

Fairness note

K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

CRM AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before you choose

K3’s second-place finish puts pressure on assumptions about which model can run a business task best. The top two scores were close, but the experiment also shows why a league table alone is not enough: models can reach the same diagnosis and still differ on whether they close the deal. Picking a model without testing it against your own work is a bet.

See the full results and plain-language findings at Firmulate’s benchmark page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Polarization and Color Perception

I’m exploring how polarization impacts color perception and why understanding this phenomenon can transform the way you see your environment.

Blue Light Hazard and Eye Health

The threat of blue light to eye health and sleep quality is real—discover simple steps to protect yourself and stay healthier in a digital world.

Achromatic Colors: Black, White, Gray

Meta description: “Master the art of using achromatic colors—black, white, and gray—to elevate your designs with timeless elegance and subtle sophistication.

Where to Find the Colors Your Screen Can’t Show You

Exploring the limits of digital color reproduction and how to experience colors beyond screen capabilities in the real world.