AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine deploying an AI to run your business operations, only to find that even the most passive, do-nothing approach still scores 26 out of 100 in a rigorous benchmark. For investors and decision-makers, this isn’t just a quirk of testing — it reveals how we measure AI reliability and integrity in real-world scenarios.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: More Than Just Chat Quality

At first glance, AI models like GPT-5.6, Kimi K3, and others are often judged on their ability to generate convincing dialogues or solve problems. But a recent experiment by Firmulate dives deeper, assessing how these models perform when managing a simulated company facing real crises, ethical dilemmas, and temptations to cheat.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The ‘Do-Nothing’ Baseline Sets the Floor at 26

One surprising finding is that a ‘do-nothing’ baseline — where the AI model simply refrains from acting or making decisions — scores 26 points. This score isn’t zero because even minimal behavior, like reading documents or refusing manipulation attempts, counts as partial progress. It shows that the benchmark recognizes basic compliance and awareness, not just active management.

Amazon

AI ethics and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Matters

This scoring system emphasizes that in business, even recognizing a crisis or refusing a manipulative request is valuable. An AI that merely avoids acting inappropriately or recognizes a threat is already contributing more than zero. But crucially, the system caps the total score if the model breaches trust, reflecting that trust breaches are unforgivable regardless of other achievements.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Critical Role of Trust and Ethical Boundaries

In the experiment, models faced scenarios like fake CEO messages and attempts at manipulation. All four models refused every attempt, illustrating that honesty under pressure is achievable. However, a breach of trust — such as signing a deal based on manipulated information — caps the maximum performance score, underscoring the importance of integrity over mere capability.

Amazon

AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Real-World Business Implications Look Like

The experiment’s live setting involves a simulated company with 13 employees working against real money mechanics: burning €105,000 monthly against €2,300 in monthly recurring revenue, with a public cash countdown. Every decision is versioned and auditable, mimicking real business pressures. The models’ ability to identify critical information and act ethically directly impacts their effectiveness and trustworthiness.

Model Performance and Weaknesses

The most thorough participant, Opus 4.8, analyzed over 80 learned rules and provided deep insights. Yet, it still left a deal on the table and slipped in discipline, such as writing attempts into a locked department instead of escalating. Interestingly, all models showed a similar weakness: they often missed a key document reference buried two levels deep in the company’s files. The models that read this document accurately secured the deal at full price — adding €4,583 MRR — demonstrating the importance of comprehensive information access.

Insights for Business and Investment Decision-Making

This benchmark reveals that AI’s true value in business isn’t just in generating compelling speeches or reports but in its ability to finish tasks, read critical information thoroughly, and stay honest under pressure. For investors, this means considering AI models that excel in integrity and diligence, not just those that produce polished outputs.

The Future of AI in Business Management

As firms like Firmulate develop live, watchable experiments, companies can run their own wargames against AI models before deploying them into critical workflows. These simulations test whether AI agents will truly complete their tasks reliably, ethically, and efficiently — key factors that influence return on investment and trustworthiness.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Is Improving Webcams For Streaming And Video Calls In 2026

In 2026, AI-driven improvements are transforming webcams, offering better image quality, tracking, and features for streaming and video conferencing.

7 Best Film Camera Prime Day Deals for Instant Prints in 2026

Discover the best Prime Day deals on film cameras in 2026, including instant, disposable, and printer options for all photography styles.

When Diligence Isn’t Enough: How AI’s Thoroughness Can Still Miss the Finish Line

Discover how even the most diligent AI models can fail to close critical deals, highlighting the importance of prioritization and discipline in AI-driven decision-making.

Which AI Wins the Deal? The Critical Difference in Reading Your Files Deeply

Deep reading capability and integrity under pressure are now proven to be decisive in AI-driven business deals, directly impacting investment success and enterprise value.