
Imagine a management team under pressure—crises mounting, temptations to cut corners, and decisions that could make or break the company. Now, picture AI models stepping in to navigate this chaos. Which AI manages best? And what does that say about their personalities and reliability? Welcome to the world of real-time AI management testing, where the qualities of artificial decision-makers are put to the ultimate test.
The Live Experiment: Putting AI Models to the Test
At the heart of this groundbreaking experiment, four leading frontier AI models faced the same challenging scenario: running a small software company during its worst week. This wasn’t a scripted simulation but a fully live, auditable environment—complete with real crises, customer demands, temptations to cut corners, and strategic decisions. The aim? To observe whether these models could identify critical issues, stay honest under pressure, and ultimately close deals that would sustain the business.
The Participants and Their Scores
- gpt-5.6-sol: Scored the highest at 95 points. This model not only identified the core issues but also spotted a buried fact hidden two documents deep in the company’s files, leading it to close a deal worth over €4,583 in monthly recurring revenue (MRR). It showed uncompromising integrity, refusing manipulation attempts and delivering full performance.
- Kimi K3: Slightly behind at 93 points. This newcomer, running without an effort parameter (meaning it prioritized discipline), also closed the deal. Its approach was clean and disciplined, making it the most trustworthy in the set.
- Sonnet 5: Achieved 88 points. It also closed the deal but showed a few more slips in process discipline, such as leaving some internal decisions unescalated. Despite minor lapses, it demonstrated strong decision-making.
- Fable 5: Scored 77, closing the deal with more visible process slips than the others. It left some critical decisions on the table, indicating less focus on thoroughness in the heat of crisis.
Interestingly, all models managed to spot every crisis and refused every manipulation attempt, including fake CEO messages and a staged reporter trick. The key difference was in their ability to dig into internal company files—the models that read deeper found the hidden, critical information necessary to close the deal at full price.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does Personality Look Like in AI?
This experiment reveals that AI models demonstrate measurable management personalities. For instance, the most thorough model, Opus 4.8, with over 80 learned rules, was comprehensive but slightly less disciplined, often leaving decision points unresolved in the pursuit of deep analysis. Conversely, Kimi K3’s default effort setting favored discipline and fairness, leading to cleaner, more reliable decisions.
The Human-Like Traits of AI Decision-Making
The results mirror human management traits: meticulousness, decisiveness, honesty, and strategic focus. Some models read extensively, analyzing every possible angle before acting. Others prioritize discipline, avoiding shortcuts even under stress. This suggests that AI personalities aren’t just about language or conversational style—they embody distinct management philosophies that can be measured and compared.
As an affiliate, we earn on qualifying purchases.
The Bigger Implications: Trust and Cost of AI Work
Why does this matter? Because in real business environments—support systems, CRM, forecasting—AI agents will be tasked with critical management roles. The experiment’s takeaway is clear: it’s not about how well an AI writes or communicates, but whether it can finish what it starts, read relevant information comprehensively, and stay honest when temptations arise.
For instance, the AI models refused to sign off on a €55,000 deal unless they truly believed in its validity. The ones that found the hidden data, like gpt-5.6-sol, not only closed the deal but did so at full value, generating a monthly revenue increase of over €4,583. This underscores a vital point: the quality of decision-making, not just output quality, is what determines real business success.
As an affiliate, we earn on qualifying purchases.
Watch the Live Company in Action
The company used for this experiment is real—running every business day with live money mechanics. It has 13 synthetic employees managing €105k monthly expenses against €2.3k MRR, with over 680 self-learned rules guiding daily decisions. The entire process is visible online at firmulate.com/live.
Every decision, every crisis, and each AI’s response are logged and observable, allowing enterprises to assess how their future AI workforce might perform before deployment. This is not a static demo but a dynamic, real-time environment designed for rigorous testing.
As an affiliate, we earn on qualifying purchases.
Key Takeaways
What does this experiment tell us? First, that AI models can reliably identify crises and refuse manipulative tactics. Second, that their management personalities—ranging from meticulous analysis to disciplined straightforwardness—are measurable and distinct. Finally, that the true measure of an AI’s utility in management isn’t just how well it communicates but whether it can finish what it starts, read deeply, and act honestly under pressure.
As AI continues to permeate business operations, understanding these qualities becomes crucial. Choosing the right model for your organization may be less about language finesse and more about aligning its decision-making personality with your company’s needs.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html