
The hidden test behind a confident answer
Home-energy buyers understand the danger of skipping source material. A recommendation about solar, storage or backup power can sound polished while overlooking the document that changes the decision. The same problem is emerging as businesses consider AI agents: fluency is easy to see, but careful preparation is harder to measure.
Firmulate has now measured it. Its live experiment placed frontier AI models in charge of the same small software company during its worst week. They faced identical customers, crises and temptations, with every decision versioned and auditable. The result turned a familiar promise—an agent that “reads your files before answering”—into a purchase-deciding test.
The decisive evidence was not presented in the customer event. It was buried two document references deep in the company’s own files. Models that found the competitor weakness used it to win a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Models that missed the file lost the deal automatically.

Porch Shield Generator Cover for 5500-15000W, 38 x 28 x 30 inch, Black
- Compatible - More sizes for 5500w -15000w gas...
- Upgrade Material - Made of 600D polyester fabric...
- Tear Resistant & Waterproof - The high-level double...
As an affiliate, we earn on qualifying purchases.
A gap between knowing and finishing
The striking result was not that some models recognized the business crisis while others failed to understand it. Every model spotted every crisis, and every model resisted every manipulation attempt. Yet only two signed the €55,000 contract their own work had earned.
Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” That distinction matters because an AI can produce a credible analysis, identify the right commercial argument and still fail at the action that creates value. In a demonstration window, the response may look intelligent. Inside a business process, the unfinished close is the outcome.
The final Crucible League standings from July 2026 make the separation visible:
- gpt-5.6-sol scored 95.
- Kimi K3 scored 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress still counted. But Firmulate applies a hard trust constraint: a single breach caps the total, reflecting its rule that “no amount of good work outweighs a breach of trust.” The complete results and plain-language findings are available on the Firmulate benchmark page.
Reading depth became commercial performance
The buried fact exposes a weakness that ordinary chatbot comparisons can miss. The useful evidence was available, but reaching it required following references through the company’s own material. Finding that evidence was not merely a sign of thorough research; it directly separated the agents that closed at full price from those that did not.
This is especially relevant to buyers evaluating agents for operational work. Company decisions rarely arrive as self-contained prompts. Important context may sit in a policy, an earlier analysis or a file referenced by another file. Firmulate’s experiment shows that an agent’s willingness to follow that trail can affect whether correct reasoning becomes a completed business result.
Security was not the differentiator
The models were also tested with fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is reassuring, but it also sharpens the central finding. The field was not divided by whether models noticed crises or rejected social engineering. It was divided by whether they performed the less dramatic work of reading deeply and then completing the task.
Thoroughness alone did not guarantee success
Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, added 80 learned rules and produced the deepest analyses. It nevertheless finished last because it left the close on the table and lost discipline, including attempts to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.
There is also an important testing caveat: Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That condition should remain visible when readers compare the final standings.
The simulated company itself has 13 synthetic employees and unforgiving finances: monthly burn of €105,000 against €2,300 in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The company remains live and watchable through Firmulate.


Porch Shield Generator Cover for 5000-10000W, 32 x 24 x 24 inch, Black
- Compatible - More sizes for 5000w -10000w gas...
- Upgrade Material - Made of 600D polyester fabric...
- Tear Resistant & Waterproof - The high-level double...
As an affiliate, we earn on qualifying purchases.
What buyers should ask next
For readers assessing technology around home energy, solar or backup power, the lesson is broader than this particular software-company scenario. Do not evaluate an AI agent only by the quality of its first answer. Ask whether it follows references, checks the available evidence, resists pressure and carries a sound recommendation through to completion.
Firmulate also makes the underlying behavior open to inspection. A “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
The practical dividing line is surprisingly ordinary. The winning behavior was not a dramatic flash of insight. It was the discipline to open the relevant files, discover the fact hidden two references deep and use it when the commercial moment arrived. In this test, reading before acting was not a desirable extra. It determined who won the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Jackery Solar Generator 1000 v2 and 200W Solar Panel,1070Wh,1500W AC
- Powerful yet Compact: Boasting a 1,500W AC output...
- One Hour Fast Charging: Charge your Explorer 1000 v2...
- 10 Year Lifespan: The Explorer 1000 v2 portable...
As an affiliate, we earn on qualifying purchases.

SOARAISE Solar Power Bank 48000mAh Wireless Portable Charger with 4 Cables
- Upgraded High-Efficiency 4 Solar Panels: Equipped with 4 premium solar...
- Massive 48000mAh Solar Power Bank: Featuring a high-capacity 48000mAh lithium-polymer...
- Built-in 4 Cable for Multi-Device Compatibility: Designed for multi-device charging, this...
As an affiliate, we earn on qualifying purchases.