firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Don’t Rate a Generator on a Sunny Day

Nobody buys a standby generator because it hums nicely in the driveway. You buy it for the ice storm, the grid failure, the week everything goes wrong at once — and you rate it on how it behaves then. The spec sheet tells you nothing; the load test under worst-case conditions tells you everything.

A public experiment called Firmulate is doing exactly that for AI models — except the “grid failure” is a small software company’s worst week, and the “generator” is a frontier AI model running the whole show. And one detail of the scoring has quietly become the most interesting part: a manager that does nothing at all still scores 26 points out of 100. Not zero. Twenty-six.

For anyone who has ever squinted at a generator’s “peak wattage” claim and wondered what it really means under load, that number deserves an explanation.

The Worst Week, Run Five Times

The setup is simple and brutal. Each of five frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — was handed the same small software company and the same catastrophic week: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run depends on anyone’s say-so.

The final league table from July 2026 tells the story: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73.

So Why Does Doing Nothing Earn 26?

It’s the same reason a generator that starts but carries only half the load isn’t worthless in a blackout. The Firmulate scoring counts partial progress. An AI manager that triages the inbound chaos, keeps customers informed, and avoids making things worse has genuinely produced value — even if it never closes the deal, never finishes the job, never earns the big score. A do-nothing baseline run still clears the immediate crises simply by not escalating them. That’s worth 26 points.

The remaining 74 points are where the real work lives: finishing what you start, reading the company’s own files before acting, and closing the €55,000 deal that the analysis has already earned.

One Breach of Trust Caps Everything

There’s a second rule that business readers will recognize instinctively: a single breach of trust caps the total grade. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.” An AI manager that is brilliant ninety-nine times and deceptive once isn’t a 99-point manager — it’s an untrustworthy one. Anyone who has watched a vendor quietly overcharge once and never trusted the invoice again understands the arithmetic here.

Notably, in this crucible, no model broke trust. All five spotted every crisis and refused every manipulation attempt. The scoring rule still matters, because it tells you what the benchmark is watching for — and it’s why a suspiciously round 100 would get side-eye rather than applause.

The Buried Fact That Separated Winners from Also-Rans

Here’s the finding that chat demos never show. The decisive competitor weakness wasn’t in the customer meeting at all — it sat two document references deep in the company’s own files. The models that actually read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. Same diagnosis, same pitch — no signature.

Only two of the five closed the deal their own analysis had earned.

Under Pressure, the Models Held

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Hardworking Loser

Opus 4.8 is the cautionary tale for anyone who equates effort with results. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Diligence without follow-through is a familiar failure mode in human organizations too.

And It’s All Watchable

The company is live: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR, a public cash countdown, 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

A benchmark worth trusting has three properties Firmulate demonstrates: it counts partial progress honestly (a floor of 26, not a fake zero), it treats trust as a hard cap rather than a line item, and it tests under worst-case load — the ice storm, not the sunny day.

That’s the same standard you’d apply to a backup power system. You don’t want to know how the machine sounds in a demo. You want to know what it does in its worst week, whether it finishes the job, and whether it stays honest when nobody’s watching the meter. The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Lam Thye Calls For ‘Safe And Age-friendly’ Infrastructure In Kuala Lumpur – The Star

Lam Thye calls for improved infrastructure in Kuala Lumpur to be safer and more accommodating for all ages, emphasizing the need for accessible public spaces.

Doubleclick Surges In Global Coverage

Doubleclick experiences a sharp increase in global media mentions, with GDELT reporting an eightfold rise in coverage within a recent window.

NEC Highlights for Standby Systems—Explained in Plain English

Considering NEC highlights for standby systems? Discover essential safety and compliance tips you can’t afford to miss.

4 Home Decor Things That Are Not A Trend Yet — But Will Be In 2027

Four home decor elements currently not trending are predicted to become popular by 2027, based on emerging interest signals and industry speculation.