AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine relying on an AI to steer your small business through its toughest week — managing crises, making trustworthy decisions, and closing deals. The latest experiment from Firmulate reveals which AI models are ready for prime time and which still stumble under pressure. For investors and business leaders, understanding how these models perform in real-world scenarios could be the key to smarter, more reliable automation.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI in the Wild

In a groundbreaking live experiment, four frontier AI models were tasked with running a small software company during its worst week — facing the same crises, customer demands, and temptations to cut corners. Every decision made was recorded and auditable, designed to mimic the real pressures a business faces daily. The goal? See which AI can detect hidden risks, resist manipulation, and ultimately close a lucrative deal.

Amazon

AI business automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: The Leaders and the Laggards

All four AI models identified every crisis and refused manipulation attempts, demonstrating a baseline of integrity and situational awareness. However, their ability to convert diagnosis into action varied significantly.

  • The top performer, gpt-5.6-sol, scored 95 out of 100, found the buried security issue, and successfully closed the €55,000 deal, adding €4,583 in monthly recurring revenue (MRR).
  • Moonshot’s Kimi K3 was a close second at 93, achieving the same deal, with the cleanest discipline and no deviations. It even found a buried document reference that clinched the deal at full price.
  • Sonnet 5 scored 88, closing the deal but with minor process slips, while Fable 5 trailed at 77, also closing but leaving some potential on the table. Opus 4.8 scored 73 and demonstrated the deepest analysis but faltered on the close.

Notably, K3 ran without an effort parameter (the default API setting), whereas others operated under high effort modes, making K3’s performance even more impressive given equal conditions.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Sets K3 Apart?

Beyond the scores, Kimi K3’s success lay in its ability to read deeply into the company’s own files — two document references deep — and uncover critical information unnoticed by other models. This ability to dig beneath surface data was decisive in winning the deal at full price, translating into a real revenue boost.

Amazon

AI cybersecurity risk detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Manipulation and Social Engineering

All models demonstrated resilience against social engineering attempts, including fake CEO messages and staged reporter tricks. K3 explicitly reasoned that requests should be treated with suspicion, reflecting a cautious, disciplined approach. This is crucial for real businesses where trust and integrity are paramount.

Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Test: A Live Company

The experiment was conducted on a simulated but fully operational company, with 13 synthetic employees, real money mechanics, and over 680 self-learned rules. The company burns €105,000 monthly against €2,300 MRR, illustrating the high stakes of automation in actual business settings. The live scenario is accessible at firmulate.com/live.

The Lessons for Business Leaders and Investors

While AI chat demos often showcase impressive language skills, this experiment proves that real-world management depends on more: reading comprehension, integrity under pressure, and the ability to execute decisions. The league table shows a clear gap between models that can identify buried risks and those that cannot, emphasizing the importance of choosing AI tools that can finish what they start.

For investors considering AI-driven automation, these findings suggest that performance scores matter — but so does the AI’s discipline and depth in analysis. A model that detects hidden risks and resists manipulation is more likely to prevent costly mistakes and secure revenue, making it a smarter bet for future-proofing your business.

Conclusion: The Future of AI in Business Management

The leaderboard from this live experiment underscores a vital point: the AI league is still open, and choosing the right model isn’t just about raw scores. It’s about understanding how these systems perform in real crisis scenarios, reading deeply into data, and maintaining discipline under pressure. For investors and business owners, the takeaway is clear — test your AI tools thoroughly before trusting them with your company’s future.

To see the models in action and explore the full results, visit firmulate.com/benchmarks.html.

Fairness Note

K3 was tested without an effort parameter (the default API setting), while the other models operated at xhigh effort. This difference is noted to ensure transparent comparison of performance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI In Home Cinema: 9 Projectors Leading The Way In 2026

Nine top projectors in 2026 showcase AI integration for home cinemas, emphasizing brightness, image quality, and smart features. Key developments and future outlook.

Which AI Wins the Deal? The Critical Difference in Reading Your Files Deeply

Deep reading capability and integrity under pressure are now proven to be decisive in AI-driven business deals, directly impacting investment success and enterprise value.

AI Models Pass the Test: Maintaining Integrity Under Pressure in Business Simulations

In a groundbreaking experiment, five top AI models faced simulated crises and social engineering tests, with all refusing manipulation—highlighting AI’s potential for trustworthy decision-making.

Couples’ Ultimate Guide To AI Wedding Planning SaaS

A new AI-driven wedding planning SaaS aims to help engaged couples manage their weddings without professional planners, testing in a pilot phase.