
A backup generator earns trust when the grid fails, not when the lights are already on. AI tools face a similar test: can they make sound decisions when a business hits its worst week? A live Firmulate experiment put five frontier models in charge of the same small software company. Moonshot’s Kimi K3 finished second, ahead of three Western models.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate’s Crucible is a watchable experiment in AI management. Each model ran the same small software company through the same customers, crises and temptations. Every workday is versioned and auditable, and the company runs with real money mechanics. Its 13 synthetic employees face a burn of €105,000 a month against €2,300 in monthly recurring revenue.
The final league, from July 2026, puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. A do-nothing baseline scored 26. The result makes the field look open: K3 beat three of the four Western frontier models in this contest.
The difference was follow-through
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor’s offer was buried two document references deep in the company’s files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue.
K3 found the buried security needle, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. The experiment’s broader lesson is that sound analysis does not guarantee action: “Same diagnosis, same pitch — no signature.”
The bait included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as results
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.
For businesses weighing AI agents, the point is practical: polished answers and careful analysis are not the whole job. A model may need to read company records, protect customers and complete a transaction under pressure. Firmulate argues for testing those behaviors in a company-like setting before trusting a model with work. The company has accumulated more than 680 self-learned playbook rules, and its live run is available at firmulate.com.
The results come with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, so readers can try to guess which model made each choice. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Full results and findings are at Firmulate’s benchmark page.

Test before handing over the controls
For home energy, a system’s performance under ordinary conditions says little about what happens in an outage. Firmulate’s experiment suggests a comparable standard for AI at work: put models through realistic pressure, then judge whether they finish the job and respect boundaries. Kimi K3’s second-place result shows that the leader board can shift, making model choice without your own test a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
