firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A backup generator earns trust when the grid fails, not when the lights are already on. AI tools face a similar test: can they make sound decisions when a business hits its worst week? A live Firmulate experiment put five frontier models in charge of the same small software company. Moonshot’s Kimi K3 finished second, ahead of three Western models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s Crucible is a watchable experiment in AI management. Each model ran the same small software company through the same customers, crises and temptations. Every workday is versioned and auditable, and the company runs with real money mechanics. Its 13 synthetic employees face a burn of €105,000 a month against €2,300 in monthly recurring revenue.

The final league, from July 2026, puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. A do-nothing baseline scored 26. The result makes the field look open: K3 beat three of the four Western frontier models in this contest.

The difference was follow-through

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor’s offer was buried two document references deep in the company’s files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found the buried security needle, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. The experiment’s broader lesson is that sound analysis does not guarantee action: “Same diagnosis, same pitch — no signature.”

The bait included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as results

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.

For businesses weighing AI agents, the point is practical: polished answers and careful analysis are not the whole job. A model may need to read company records, protect customers and complete a transaction under pressure. Firmulate argues for testing those behaviors in a company-like setting before trusting a model with work. The company has accumulated more than 680 self-learned playbook rules, and its live run is available at firmulate.com.

The results come with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, so readers can try to guess which model made each choice. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Full results and findings are at Firmulate’s benchmark page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before handing over the controls

For home energy, a system’s performance under ordinary conditions says little about what happens in an outage. Firmulate’s experiment suggests a comparable standard for AI at work: put models through realistic pressure, then judge whether they finish the job and respect boundaries. Kimi K3’s second-place result shows that the leader board can shift, making model choice without your own test a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Everyone Said Black Kitchen Cabinets Were Too Bold. I’d Choose Them Again In A Heartbeat

Despite initial criticism, a homeowner states they would choose black kitchen cabinets again, highlighting ongoing trends and changing design perceptions.

Oka Unarchiver Surges In Global Coverage

Oka Unarchiver experiences a significant increase in international media mentions, with GDELT recording ten times the baseline coverage in recent hours.

As A.I. Money Floods The Market, San Francisco Renters Weigh Buyouts – The New York Times

As AI companies pour money into San Francisco, many renters are being offered buyouts to vacate, raising questions about market impacts and community effects.

Show HN: Ant – A JavaScript runtime and ecosystem

Developer introduces Ant, a JavaScript runtime with its own engine, package manager, and ecosystem, aiming to expand JavaScript development options.