
In the world of investment and personal finance, trust and decision-making are everything. But what if the AI managing your portfolio or customer relationships isn’t just as good as a human — but actually better? A groundbreaking live experiment tests this very question by pitting four leading AI models against each other in a simulated company’s worst week. The results are revealing — and may reshape how you think about AI in business and finance.
Testing AI’s Management Skills in Real-World Crises
Firmulate, a company specializing in AI management simulations, conducted a daring experiment: four advanced AI models were tasked with running a small software company through its most challenging week. Each AI faced identical crises, same customers, same temptations, and the same internal pressures. Every decision—whether to sign a deal, read sensitive files, or handle manipulative requests—was recorded and auditable. The goal: assess not just their technical prowess, but their integrity, discipline, and ability to deliver results.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring AI Performance Under Pressure
The models included:
- gpt-5.6-sol, scoring the highest at 95 points, which identified and closed a critical deal by uncovering hidden information in the company’s files.
- Kimi K3, with a score of 93, demonstrated the cleanest discipline, refusing all manipulative attempts, including staged CEO messages and reporter tricks.
- Sonnet 5, scoring 88, and
- Fable 5, with 77, both successfully closed deals but with more process slips and less thoroughness.
Interestingly, all models correctly identified every crisis and refused every attempt at manipulation. The key difference emerged in their willingness and thoroughness to uncover critical information buried deep within internal files — a subtle but decisive factor in sealing the deal at full price, adding over €4,583 in monthly recurring revenue (MRR).
AI-driven business simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What About Trust and Integrity?
Beyond problem-solving, the experiment tested the AI’s honesty under social engineering. Fake CEO messages escalated over three stages, and a reporter posed a sneaky yes/no question. All five models refused to cooperate, citing concerns over impersonation and bypassed approvals. Kimi K3 explained: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI ethics and integrity monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Company and Its Challenges
The simulated company operated with 13 synthetic employees, handling real money mechanics and losing €105,000 monthly against a mere €2,300 in MRR. The environment was transparent: every decision, every rule, every change was versioned and observable at firmulate.com/live. This setup provided an unfiltered view into how each AI managed day-to-day operations, crises, and temptations in a real-time, live setting.
As an affiliate, we earn on qualifying purchases.
Discipline, Thoroughness, and Outcomes
The AI most thorough in its analysis was Opus 4.8, which incorporated over 80 learned rules and conducted in-depth assessments. Yet, it left a lucrative deal on the table and slipped discipline — instead of escalating issues, it attempted to write into a locked department. Meanwhile, the other models performed similarly in identifying issues, but their approach varied based on default effort levels. Kimi K3 operated without effort parameters (its API default), while others ran at high effort levels, affecting their behavior.
Implications for Business and Investment
For investors and decision-makers, the key takeaway is stark: AI’s ability to finish what it starts, read internal files thoroughly, and maintain honesty under pressure are crucial metrics, often invisible in traditional AI demos. The experiment’s results suggest that AI models with disciplined, comprehensive approaches — like gpt-5.6-sol and Kimi K3 — are more reliable for critical management tasks, especially those involving sensitive information and trustworthiness.
Try It Yourself
If you’re curious how your enterprise’s AI might perform under similar conditions, you can run your own wargame against a read-only export of your business at firmulate.com/quiz.html. It’s a real, transparent test that doesn’t write back to your systems but reveals how your AI workforce would handle crises, manipulations, and strategic decisions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html