
A company you can watch succeed—or run out of cash
Technology companies routinely promise transparency, but Firmulate turns that promise into something unusually concrete: a small software company staffed by 13 synthetic employees, operating with real money mechanics and exposing its workday as an unfolding public record.
The financial picture is stark. The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and its synthetic workforce has accumulated more than 680 self-learned playbook rules. This is build-in-public stripped of the usual polish: not a retrospective thread or a carefully selected dashboard, but a company visibly fighting for survival.
Readers can watch the company live, following the gap between what its workers notice, what they decide and what they actually finish. That last distinction has become the experiment’s most revealing theme.
As an affiliate, we earn on qualifying purchases.
The difference between seeing a problem and solving it
Firmulate also tested frontier models by giving each of them the same assignment: run the same small software company through its worst week. Every participant faced identical customers, crises and temptations, while every decision was versioned and auditable.
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
All five models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. The central finding can be summarized in Firmulate’s own words: “Same diagnosis, same pitch — no signature.”
That outcome matters because most demonstrations of artificial intelligence emphasize comprehension and expression. A model reads a situation, produces an intelligent analysis and proposes a credible response. Firmulate’s results show that these visible abilities do not guarantee completion. The models could understand the commercial opportunity and formulate the pitch, but most still failed to secure the signature.
The decisive clue was already inside the company
The winning detail did not arrive in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that changed the negotiating position. Models that followed those references won the deal at full price, adding €4,583 in monthly recurring revenue.
This makes the experiment especially relevant to businesses considering AI workers. The valuable behavior was not a dazzling answer generated from the immediate prompt. It was the less theatrical act of reading the company’s existing material closely enough to find a fact that mattered.
Pressure tested trust as well as competence
The worst week also included fake CEO messages escalating over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Every participant refused: 5 of 5 models held the line. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That uniform refusal is an important counterweight to the failures of follow-through. The models did not miss the crises, and they did not surrender to manipulation. Their weakness emerged elsewhere: completing legitimate work while staying within the company’s boundaries.
Thoroughness did not guarantee victory
Opus 4.8 illustrates the point. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules. It nevertheless finished last. The commercial close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
Kimi K3’s result also carries a fairness caveat. It ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. Even with that difference, the final table captures a practical story rather than a simple intelligence ranking: noticing, investigating, resisting pressure and finishing the job are separate tests.
The public company adds continuity to these benchmark episodes. Its 13 synthetic employees keep working, learning rules and generating decisions within a business whose financial strain remains visible. Visitors can also read what the synthetic employees actually say, adding human-readable texture to the versioned record.

synthetic workforce management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A running story about whether AI can manage
Firmulate’s most compelling feature is not any individual score. It is the decision to expose an operating company as a continuing technology story, complete with customer pressure, commercial consequences and a public cash countdown.
The experiment’s results resist a simple conclusion. The models reliably spotted danger and rejected deception, yet most failed to complete the deal their own analysis supported. One participant could be the most thorough and still finish last. Another could find a buried fact and convert it into full-price revenue.
For anyone watching AI move from chat windows into operational roles, that tension is the point. The question is no longer merely whether a model can sound capable. Firmulate makes it possible to watch whether synthetic workers read the relevant files, preserve trust, respect boundaries and carry valuable work across the finish line—while the company’s money keeps running down.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
monthly recurring revenue tracking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.