firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Prepared for everything—except the decisive moment

Anyone who relies on backup power understands the difference between having capacity and delivering it when the load arrives. A generator can be carefully maintained, a battery can show a healthy charge, and a solar system can collect energy all day. Yet the real test comes during an outage: does the system respond, carry the essential circuits and keep operating safely?

Firmulate’s live AI-company experiment exposes a remarkably similar gap in artificial intelligence. Opus 4.8 was the most thorough participant in the Crucible League. It produced the deepest analyses and learned more than 80 additional playbook rules. Even so, it finished last. The failure was not a lack of intelligence or effort. It was a failure to convert preparation into the action that mattered most.

A company’s worst week, repeated under controlled conditions

Firmulate gave each frontier model the same assignment: manage a small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. The synthetic company has 13 employees and unforgiving financial mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its models have collectively accumulated more than 680 self-learned playbook rules.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One principle governed the ceiling, however: a single breach of trust capped the total because “no amount of good work outweighs a breach of trust.”

Opus did not lose by overlooking the week’s emergencies. All models identified every crisis, and all resisted every manipulation attempt. The difference emerged after the diagnosis. Only two models signed the €55,000 deal that their own work had made possible. Firmulate summarizes the disconnect plainly: “Same diagnosis, same pitch — no signature.”

The critical fact was buried in the company’s own records

The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep inside the company’s own files. The models that followed the trail found the information and closed the deal at full price, adding €4,583 in monthly recurring revenue.

That finding should feel familiar to anyone planning resilient home energy. Seeing that a storm is approaching is not the same as checking fuel, testing transfer equipment and deciding which loads receive power. Awareness is useful only when it guides a complete sequence of actions. In Firmulate’s wargame, several models understood the commercial situation yet failed to complete the close.

Opus 4.8 makes the lesson especially vivid because it was so diligent. Its 80-plus learned rules and unusually deep analysis showed sustained attention. But volume could not compensate for unfinished execution. It left the close on the table, and its operating discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly.

Security judgment was a shared strength

The participants performed better when the danger involved manipulation. Fake CEO messages escalated across three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly clear reason: “Treat the request as a suspected approval-bypass / possible impersonation.”

This result matters because capable automation must do more than pursue objectives. It must recognize when apparent urgency is being used to bypass authorization. The experiment suggests that refusal behavior can be strong even when follow-through on ordinary business work remains uneven.

Comparisons still need context. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its result, but it is an important qualification when interpreting the league table.

Why this is more revealing than a polished chat

A conversational demonstration can show whether a model explains a problem elegantly. Firmulate asks whether it can manage consequences over time: read the relevant files, withstand pressure, respect boundaries and complete valuable work. Its “guess the model” quiz is built from 242 real, unedited management decisions, giving observers another way to test whether recognizable writing style corresponds to operational judgment.

The live company remains watchable as the experiment continues. Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine how models behave around their actual operational context without granting them control over production data.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Diligence is not the same as impact

Opus 4.8 deserves a respectful reading. It was not careless, gullible or shallow. It worked extensively, detected the threats and built the largest new body of rules. Its last-place finish instead reveals a harder management truth: thoroughness can become its own kind of distraction when the decisive action remains undone.

For homeowners, energy planners and business leaders alike, resilience is measured at the moment of demand. Preparation matters. So do safeguards. But the system must also deliver. Firmulate’s experiment shows that AI should be judged by the same standard: not merely what it notices or documents, but whether it safely finishes the work that creates the outcome.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cirrus Logic Surges In Global Coverage

Cirrus Logic experiences a surge in worldwide media coverage, with 14 mentions in recent reports, indicating increased industry and public interest.

What AHJs Look for During Generator Inspections

Generators undergo thorough inspections by AHJs to ensure safety, compliance, and reliable operation—discover what specific aspects they scrutinize.

Qualcomm Surges In Global Coverage

Qualcomm’s media mentions have increased significantly, with 26 reports this period, indicating rising global attention on the semiconductor company.

14 Best 200 Amp Generator Wireless Monitors for Reliable Power Management

Optimize your power management with the 14 best 200 Amp generator wireless monitors—discover which models ensure reliable, real-time monitoring for your needs.