
Imagine a company with no human employees, yet it’s running an entire business—facing crises, making decisions, and even losing money every day. This isn’t fiction; it’s a live experiment revealing how AI can govern a real company under pressure, all in full view of the public.
What Is This Live Company Experiment?
At the heart of this bold venture is Firmulate, a platform that runs what it calls a “live AI company.” Unlike traditional startups, this one operates with 13 synthetic employees—AI models that make all decisions, respond to crises, and execute tasks, just like human staff. But crucially, it’s a real company with real money mechanics, currently burning €105,000 per month against a modest €2,300 Monthly Recurring Revenue (MRR). The company’s daily operations are public and transparent, with every decision versioned and auditable, making it an open window into AI’s capabilities and limitations in business management.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI Under Extreme Conditions
Each day, four different advanced AI models are given the same challenging scenario: run a small software firm through its worst week, with the same customers, crises, and temptations to cheat or manipulate. These models are fed the company’s latest data, internal documents, and playbook rules—over 680 learned and self-adjusted protocols—allowing them to navigate complex business decisions.
AI decision-making tools for startups
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: AI’s Strengths and Weaknesses
- All four models — gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8 — detected every crisis and refused every manipulation attempt, demonstrating strong integrity and crisis recognition.
- Despite this, only two models managed to close a crucial €55,000 deal, which their own analysis justified, giving them the full increase in monthly revenue (+€4,583 MRR). The other two, despite diagnosing correctly, left opportunities unexploited, leaving potential revenue on the table.
- Hidden in the company’s internal files was a decisive piece of information that the models that read these documents won the deal—highlighting the importance of comprehensive information access for AI decision-making.
- The experiment also tested social engineering attacks—fake CEO messages, escalations, and reporter tricks—all of which were consistently rejected by the models, with Kimi K3 explicitly treating such requests as potential impersonations.
As an affiliate, we earn on qualifying purchases.
The Real Company in Action
The live setup isn’t just an AI simulation; it’s a functioning software company operating daily, with real money mechanics. Every day, the company burns €105k but earns just €2.3k, facing an urgent cash countdown. It relies on over 680 learned rules and is constantly versioned, with every decision documented for review. Watch it live to see decisions unfold and learn how AI handles crises, manipulations, and strategic opportunities.
AI enterprise automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights from the Results
The performance scores from the Crucible League provide context:
- gpt-5.6-sol led with a score of 95, successfully uncovering hidden critical information and closing the deal.
- Kimi K3, a newcomer, scored 93, also closing the deal with the cleanest discipline seen so far.
- Sonnet 5 scored 88, closing the deal but with slightly more process slips.
- Opus 4.8 scored 73, demonstrating deep analysis but falling short on closing opportunities, illustrating that thoroughness alone does not guarantee success.
This experiment is a clear demonstration: AI can recognize and respond effectively to crises, refuse manipulative tactics, and, under the right conditions, close deals and execute complex decisions. Yet, it also reveals that information access and disciplined execution remain critical—and that AI’s decision-making in a live, money-losing environment is still a work in progress.
Why This Matters for Business and Education
For anyone interested in how AI could reshape organizational management, this live experiment offers invaluable insights. It questions whether AI can truly replace human judgment, especially under stress and temptation, or if it’s better suited to assist. It also underscores the importance of transparency and build-in-public approaches, where every step in decision-making is shared openly, allowing scrutiny and improvement.
To explore the decisions yourself, take the management quiz with 242 real, unedited decisions, or run your own business scenario against a read-only export of this live company at Firmulate’s pilot tool.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html