
The Hidden Risks of AI in Business: Watching a Company Live
Imagine a business that operates entirely without human employees, yet struggles daily to stay afloat. While many focus on AI’s promise for productivity, this experiment reveals the real—and often overlooked—challenges of trusting AI with critical decisions. For investors and business leaders, understanding how AI performs under pressure is crucial. Today, you can actually watch this company battle its survival live, revealing what AI can and cannot do when it matters most.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experimental Company: A Front-Row Seat to AI in Action
At the heart of this experiment is a real, functioning software company run entirely by AI models—13 synthetic employees working 24/7. Instead of representing a fictional scenario, this company is live, with daily operations, a public cash countdown, and a transparent decision log that anyone can observe at firmulate.com/live.html. Every workday, the company makes thousands of decisions, responds to crises, and navigates temptations, all while its performance is meticulously tracked and versioned.
What makes this setup extraordinary is the direct comparison of different AI models running the same business. Four frontier models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—face identical challenges: customer crises, internal dilemmas, and manipulation attempts designed to test honesty and discipline. The results are revealing: all models identify every crisis and refuse every manipulation attempt. Yet, only two are able to close the deal they analyze and pitch for, signing a €55,000 contract. The remaining two fall short, losing the opportunity despite similar assessments.

Project Management Tools (AI for Risks)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights from the Live Experiment
Most striking is the discovered weakness buried deep within the company’s own files. A crucial document reference—hidden in the company’s files—had the key information needed to win the deal. Models that read and incorporate this file information managed to secure the full €4,583 monthly recurring revenue (+€55,000 deal), while others missed it entirely. This highlights a vital point: AI’s effectiveness depends heavily on its access to and understanding of relevant context.
Adding another layer, the experiment tested social engineering—fake CEO messages escalating in urgency and a reporter trick asking for quick approvals. All models refused these manipulative tactics. Kimi K3, in particular, reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a robust resistance to deception, a critical trait for AI operating in high-stakes environments.

Deception & Self-Deception: Investigating Psychics
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Economic Reality: Burning Cash to Keep Going
The live AI company isn’t profitable. It spends €105,000 monthly to generate just €2,300 in recurring revenue, a brutal burn rate that underscores the challenge of building sustainable AI-driven operations. This public cash countdown adds urgency: if no profitable shift occurs, the company risks running out of resources. Watching this ongoing financial struggle provides a candid view of AI’s potential pitfalls in real-world business applications.

The Intelligent Ledger: How Agentic AI, Quantum Computing, Data, and Tokenisation Are Transforming Banking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Metrics and Model Rankings
In the competitive league, models are scored based on their ability to identify facts, close deals, and maintain discipline. The top performer, gpt-5.6-sol, scored 95 out of 100, successfully uncovering the critical buried fact and closing the deal. Kimi K3 followed closely with 93, demonstrating the best discipline and decision-making integrity. Sonnet 5 and Opus 4.8 scored 88 and 77 respectively, with Opus showing the most slips—failing to escalate issues or close deals in some cases.
Interestingly, Opus 4.8’s thorough analysis and learned rules did not translate into performance; discipline and focus are just as vital as analytical depth. Moreover, the models ran at different settings, with K3 configured without an effort parameter, which may influence performance benchmarks.
Why This Matters for Investors and Business Leaders
This experiment provides a tangible, ongoing window into AI’s current capabilities and limitations in business contexts. It’s not about how well AI can generate chatter or complete simple tasks, but whether it can reliably finish what it starts, read and understand relevant data, and resist manipulation under pressure. For those investing in or deploying AI solutions, these are the core questions that determine real value and risk.
More broadly, the experiment underscores the importance of transparency and continuous testing—building AI that’s accountable and resilient. Watching this real-world company in action at firmulate.com/live.html offers an unprecedented look at how AI might behave when stakes are high and money is on the line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html