firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A backup generator is tested before the lights go out. The same logic applies to AI entrusted with customer service, sales or operations: see what it does in a crisis before the crisis is real. Firmulate has been putting AI models through a watchable business experiment—and now invites companies to try the test with their own data.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, played out in public

Firmulate’s live experiment runs a synthetic software company with 13 employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its employees have learned more than 680 playbook rules. The company is synthetic; the pressure is designed to make business decisions visible.

In the final Crucible League, published in July 2026, each frontier model faced the same small company, customers, crises and temptations. The experiment’s central finding was striking: all the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch—no signature.

The league’s final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust caps the total: no amount of good work outweighs a breach of trust.

The detail that changed the deal

The deciding competitive weakness was not in the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The result points to a practical challenge for businesses: an AI agent may need to connect evidence across company records and then follow through on what it has concluded.

Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work did not guarantee a close

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that same weakness showed up in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league offers a concrete snapshot of performance under those conditions, not a claim that every model will behave identically in every company.

From watching to a company’s own pilot

The public experiment lets readers watch decisions unfold. A separate quiz uses 242 real, unedited management decisions and asks visitors to guess which model made them. Together, the live company and quiz make the benchmark tangible: the question is not only what a model says, but whether it finds the evidence, respects boundaries and completes the job.

For an enterprise, Firmulate’s proposed next step is a pilot built from a read-only export of its business. The export can represent the company’s own customers, pipeline and rules; crisis scenarios can then test how models handle pressure against that context. The intended output is a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That boundary matters. A company can examine how an AI workforce handles a simulated cash squeeze, customer problem or manipulation attempt without giving it permission to change live records. Like testing backup power before an outage, the point is to learn where the plan holds—and where it needs work—while there is still time to respond.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s experiment shows a gap between recognizing the right move and carrying it through. A pilot can test that gap against a company’s own read-only data, while keeping real systems untouched. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Devtools Must Be Open Source

Growing movement advocates for all developer tools to be open source, citing transparency and security benefits amid ongoing industry debates.

Fubo quietly raises prices. Is it still worth considering over YouTube TV?

FuboTV has quietly increased its subscription prices. This raises questions about its value compared to YouTube TV for consumers considering streaming options.

They Live Abroad, But Turned Their Tel Aviv Second Home Into A Sleek Retreat – Ynetnews

Foreign residents in Tel Aviv have renovated their second homes into stylish retreats, blending international taste with local charm.

OnePlus halts operations in USA and Europe

OnePlus announces it is stopping its business activities in the US and European markets, impacting customers and future plans.