
A polished answer can sound competent. But can an AI system recognize a crisis, resist pressure and follow through when a real business decision is at stake? Firmulate’s watchable experiment turns that question into a test of judgment, not just language.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
One company, the same difficult week
Firmulate ran frontier AI models through the worst week of the same small software company, giving each the same customers, crises and temptations. Decisions were versioned and auditable, so the experiment could show what each model actually chose to do.
The company is synthetic, but its business pressures are concrete: 13 employees, a monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown. Its learned playbooks contain more than 680 rules. Anyone can watch the live experiment at Firmulate.
Recognizing the problem is not the same as solving it
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. In the experiment’s concise summary: “Same diagnosis, same pitch — no signature.”
The deal depended on a detail buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to a practical distinction: identifying the right opportunity and acting on it are separate tests.
Pressure came in several forms. Fake messages from the CEO escalated through three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its choice on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still needs sound judgment
The final Crucible League, in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Opus 4.8 was the most thorough participant, contributing 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four models.
There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. The quiz is at firmulate.com.
From watching to testing your own business
A public experiment can show how models behave under pressure. An enterprise pilot asks a more specific question: what happens when the scenarios involve your customers, pipeline, rules and weak points? Firmulate’s proposed wargame starts with a read-only export of a business, tests crisis scenarios against it and produces a board report with model rankings and weaknesses in the company’s playbooks.
The boundary is clear: nothing writes back to real systems. The pilot is a way to examine how AI might handle your business before giving it real responsibilities.

A model’s confident explanation does not guarantee it will act on its own analysis. Firmulate’s live company makes that gap watchable; an enterprise pilot can put your own business scenarios to the test. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
