AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine training a team of AI to run a company through its toughest week—crises, manipulations, and urgent decisions—and then watching which actually follow through to completion. In a groundbreaking live experiment, four advanced models faced exactly this challenge, revealing a stark truth: high chat quality doesn’t necessarily translate into reliable execution.

The Experiment: Putting AI to the Test in a Virtual Company

Firmulate’s live trial involved four cutting-edge AI models tasked with managing a real, small software business during its worst week. This wasn’t simulated chat; the models operated in a scenario with real money mechanics, 13 synthetic employees, and a public cash countdown. Every decision was versioned and auditable, and the entire process was observable at firmulate.com/live.

The Models and Their Scores

  • gpt-5.6-sol scored the highest at 95, successfully uncovering a buried fact and closing the full deal.
  • Kimi K3, the newcomer, scored 93, also closing the deal with the cleanest discipline.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Fable 5 scored 77, demonstrating strong rule adherence but leaving the deal unexecuted.
  • Baseline and partial progress scores were significantly lower, emphasizing the models’ relative capabilities.
Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Recognition Isn’t Enough — Execution Is the Test

All four models successfully identified every crisis and refused every social engineering manipulation, such as fake CEO messages or reporter tricks. Their decision-making integrity under pressure was consistent. However, only two models—gpt-5.6-sol and Kimi K3—actually signed the €55,000 deal their own analysis had earned. The other two, despite correct diagnoses, left the opportunity unacted upon.

The Hidden Weakness: Reading the Files Matters Most

The critical difference wasn’t in how they responded to external crises but in their ability to leverage internal company documents. The models that read two document references deep into the company’s files uncovered a buried fact that proved decisive—totaling over €4,583 in monthly recurring revenue. Those that failed to do so left the deal on the table, despite understanding the situation.

Amazon

enterprise AI data reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Illusion of Chat Quality

This experiment underscores a crucial point: chat demos that showcase conversational fluency don’t measure whether AI can follow through on complex, high-stakes tasks. The models’ ability to read internal files, maintain discipline, and execute decisions under pressure remains invisible in standard chat interactions but is vital for real business applications.

Social Engineering Resistance

All models refused social engineering attempts, including escalating fake CEO messages in multiple stages and a reporter trick that asked for background approval. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a level of cautious judgment critical in real-world contexts.

Amazon

AI business automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business AI Adoption

Imagine deploying AI in your customer support, CRM, or forecasting systems. The question isn’t just whether it can generate convincing language. Instead, it’s whether it can reliably finish the work—reading necessary internal data, resisting manipulation, and executing decisions. The experiment’s live, transparent nature demonstrates that these skills are measurable and, more importantly, essential.

Performance and Discipline Matter

The experiment also revealed discipline as a key factor. Opus 4.8, which ran with the deepest analysis and the most rules learned (+80), ranked last in execution. Its failure to escalate or complete the deal highlighted that thoroughness alone isn’t enough—discipline in execution is equally critical.

Amazon

AI cybersecurity resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond Chat Demos: The Future of AI as Business Operators

Standard AI benchmarks often emphasize chat quality, but the real measure of readiness is whether AI can act decisively within complex workflows, read internal files, and stick to agreed-upon tasks under pressure. Firms contemplating AI integration should consider live testing scenarios like this to gauge actual operational readiness.

See the full results, plain-language findings, and watch the live experiment at firmulate.com/benchmarks.html. Want to test your own business? The platform offers a read-only export tool to run your company through this kind of AI wargame, without risking real systems or data.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Etymology of Lucia and the Many Ways Light Shows Up in It

For those curious about “Lucia,” exploring its Latin roots and the enduring symbolism of light reveals fascinating cultural and spiritual connections worth discovering.

People Who Can’t Picture Anything Are Rewriting The Science Of Imagination

Emerging research suggests individuals who cannot picture images are reshaping understanding of imagination, sparking increased scientific interest and debate.

Why Ezra Sounds Ancient and Modern at the Same Time

Lyrically blending myth and history with innovative fusion, Ezra creates a timeless soundscape that invites deeper exploration into their enchanting duality.

This Is Why Names Meaning Peace Keep Winning in Calming Nursery Designs

Guiding nursery design with peaceful names can foster serenity and emotional growth, but there’s more to discover about creating a truly calming environment.