AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A polished AI demo can make a model look ready for the office. A tougher test is whether it can keep its head when a customer is wavering, an impostor is pushing for approval, and the right answer is buried in company files. Firmulate’s live experiment puts AI models inside a small company and lets the public watch what they do.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same company, same worst week

Firmulate ran frontier models through the same software company, customers, crises and temptations. The experiment’s final Crucible League, in July 2026, put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models missed the crises. Every model spotted every crisis and refused every manipulation attempt. The gap appeared at the finish: only two signed a €55,000 deal that their own analysis had earned. Firmulate sums up the finding as “Same diagnosis, same pitch — no signature.”

The clue was in the files

The deal hinged on a competitor weakness hidden two document references deep in the company’s files. It was not included in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a practical point for businesses assessing AI: recognizing a problem is not the same as assembling the evidence, making a decision and carrying it through.

Trust faced its own test. Fake CEO messages escalated through three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet performance was not simply a matter of how much work a model appeared to do. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped in discipline, attempting to write into a locked department instead of escalating. Firmulate says the same weakness appeared, less strongly, in all four models.

From watching to testing your own business

The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. Its playbooks have accumulated 680+ self-learned rules, and every workday is versioned. The numbers make the experiment watchable, while the decisions show how different models handle pressure. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.

For companies, the next step is a pilot using a read-only export of their own business. They can put a digital twin through crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That makes the proposal more concrete than asking whether an AI sounds capable: it gives a business a way to see how models respond to its own customers, rules and pressure points.

One caveat accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can watch the live experiment and explore the results at Firmulate.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks under pressure

Firmulate’s experiment shows why business readiness involves more than spotting trouble or refusing a suspicious request. The models also had to find evidence, respect boundaries and close a deal. Enterprise teams can test those demands against their own business using a read-only export, with no write-back to operational systems. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Colorblindness: Types and Implications

Just understanding the different types of colorblindness can reveal surprising impacts on daily life and how you might adapt.

Color Psychology: Emotions and Marketing

AIThis post was created with the assistance of artificial intelligence (AI).Color psychology…

Achromatic Colors: Black, White, Gray

Meta description: “Master the art of using achromatic colors—black, white, and gray—to elevate your designs with timeless elegance and subtle sophistication.

Blue Light Hazard and Eye Health

The threat of blue light to eye health and sleep quality is real—discover simple steps to protect yourself and stay healthier in a digital world.