
Imagine a world where you can watch artificial intelligence not just chat but actually run a business — making critical decisions, risking real money, and facing crises in live time. At Firmulate, this is no science fiction. The company has created a closed simulation where 13 AI models act as employees, managing a tiny software firm through its worst week, with every move public and auditable.
The Experiment: AI as a Business Executive
Centering on a small, virtual software company, the live experiment pits four advanced AI models against real-world challenges. Each model runs the same small business, encountering identical crises, customer demands, and ethical dilemmas — but only one consistently makes the right calls and closes a deal worth over €55,000 per month in recurring revenue.
What makes this experiment unique? It’s run in real time, every decision versioned and publicly accessible. The models are subjected to the company’s own documents, not just surface-level chat prompts, revealing how deeply they read and interpret information. For example, the model that discovered a hidden document reference in the company’s files closed the deal at full price, adding over €4,500 to monthly revenue. Meanwhile, others failed to follow the full chain of evidence.
Crises, Trust, and Ethics in Action
The experiment also tests social engineering and manipulation attempts. Fake CEO messages and reporter tricks are used to see if the AI models can recognize and refuse unethical pressure. All four models refused to sign manipulated deals, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass or impersonation.” This highlights that the models are not just generating text but actively assessing trustworthiness.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance and Lessons
Despite similar diagnoses and pitches, only two models managed to sign the deal based on their own analysis. The most thorough participant, Opus 4.8, analyzed over 80 learned rules and conducted deep assessments, yet still left the deal unclosed, illustrating that even meticulous analysis isn’t enough if discipline slips. The other models with fewer rules or default settings missed crucial details or faltered under pressure.
This experiment underscores a critical point for businesses considering AI integration: an AI’s ability to finish what it starts, read relevant information thoroughly, and remain honest under pressure is vital. It’s not just about generating convincing language but about executing tasks reliably and ethically.

Effective Social Media Marketing: The Fast Track to Stay Ahead of the Algorithms and Create AI Magic to Supercharge Your Brand and Maximize ROI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a No-Employee Company
The firm operates with 13 synthetic employees, burning €105,000 a month against a mere €2,300 in monthly recurring revenue, with a public countdown clock tracking its cash reserves. Every day is a battle for survival, and the company’s decisions are visible and auditable, creating a living story of a startup fighting to stay afloat with AI at its core.
The Deep Dive: Model Capabilities and Failures
The models are scored in a competitive league, with GPT-5.6-sol leading at 95 points, mostly for uncovering the hidden document and closing the deal. Kimi K3 follows closely at 93, demonstrating the highest discipline in refusing unethical solicitations. Sonnet 5 and Fable 5 trail behind, with 88 and 77 points respectively, showing that even well-behaved models can miss opportunities or slip in discipline.
This performance reveals a vital insight: the most capable model isn’t necessarily the one that performs the best in every aspect, but the one that balances thoroughness, discipline, and honesty under pressure.

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for the Future
For businesses deploying AI today, the question isn’t solely about how well an AI generates language or answers questions. It’s whether it can reliably execute tasks, interpret complex information, and resist manipulation. The live experiment makes this clear: models that read deeply and refuse unethical offers are more trustworthy in critical roles.
Additionally, the platform allows companies to run their own ‘wargames’ against their data, testing how AI might react in their specific contexts without risking real systems — a powerful way to prepare for wider AI adoption.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the Fight for Survival
The entire scenario is accessible at firmulate.com/live, where you can watch the AI-driven company’s daily struggles unfold. It’s a raw, unvarnished look at what building an AI-powered company entails — from crisis management to ethical decision-making, all under public scrutiny.
This experiment demonstrates that AI’s potential extends beyond language generation into real-world decision-making — but only if we understand its strengths and weaknesses in context. Watching this small company’s daily grind offers a glimpse into the future of autonomous, accountable AI in business.

In a public, real-time simulation, AI models run a virtual company through crises and ethical dilemmas — revealing that trustworthiness, thoroughness, and discipline are crucial for AI’s role in business. The experiment shows that AI can be tested outside of chat, in high-stakes situations, to ensure it performs reliably when it matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html