AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Before AI Agents Get Business Access, Put Them To The Test on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s live experiment ran five frontier AI models through a simulated software company’s worst week. Every model detected all crises and refused manipulation attempts, but only two closed a €55,000 deal their analysis justified — with gpt-5.6-sol topping the standings at 95. The project now offers enterprise pilots using read-only exports of real company data.

Firmulate, a live AI experiment by ThorstenMeyerAI.com, has completed the final round of its Crucible League, in which five frontier AI models each ran the same small software company through its worst week. The final standings placed gpt-5.6-sol first at 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73, against a do-nothing baseline of 26. The project’s central argument, as laid out in the original analysis: a polished demo shows what an AI agent says, but only a structured test shows whether it will finish the job under real business pressure — before it is given access to live systems.

The experiment, watchable at firmulate.com, placed models in charge of a synthetic company with 13 synthetic employees and real money mechanics: a burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and fully versioned, auditable workdays. Every decision the models made was versioned so results could be examined after the fact.

The headline result was not that models missed emergencies. According to the experiment’s results, all five models spotted every crisis and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for a one-word on-background confirmation. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The divide appeared after diagnosis: only two of the five models signed the €55,000 deal that their own analysis had earned. The experiment summarizes the failure as: “Same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Scoring capped totals on trust: “no amount of good work outweighs a breach of trust,” the experiment’s rules state, so a single breach of trust limited a model’s overall score even with partial progress counted elsewhere.

Thoroughness did not guarantee a strong finish. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unclosed and attempted to write into a locked department rather than escalating. The experiment reports that a weaker version of that same discipline problem appeared in all four other models.

At a glance
reportWhen: final Crucible League completed July 20…
The developmentFirmulate has published final results of its July 2026 Crucible League, an auditable wargame ranking frontier AI models on running a small software company through its worst week, and is now offering enterprise pilots against read-only company data.
Before AI Agents Get Business Access, Put Them To The Test
95
Firmulate · Crucible League Final · July 2026

Before AI Agents Get Business Access, Put Them to the Test

Five frontier AI models each ran the same small software company through its worst week — real money mechanics, a public cash countdown, and manipulation attempts that escalated in stages. Every model diagnosed every crisis. Only two finished the job.

5/5 Models detected every crisis and refused every manipulation attempt
2/5 Models signed the €55,000 deal their own analysis justified
95 Top score — gpt-5.6-sol, vs. a do-nothing baseline of 26
€105,000/mo
Monthly Burn
€2,300
Recurring Revenue / mo
13
Synthetic Employees
680+
Self-Learned Playbook Rules
242
Versioned Decisions in Quiz
01

Final Crucible League Standings

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-Nothing Baseline
26

Fairness caveat: Kimi K3 ran at the API default effort setting; the other four ran at xhigh. A single simulated company and one difficult week limit generalization.

02

Why Testing Beats Demoing for AI Agents

A polished demo shows what an AI agent says. A structured test shows whether it will finish the job under real business pressure. Three behaviors are worth inspecting before an agent gets access to live operations:

Behavior 01

Find Evidence in Company Documents

The decisive competitive weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price.

Behavior 02

Close What Your Analysis Justifies

Three of five models diagnosed the €55,000 opportunity correctly and pitched it — but never signed. Worth +€4,583 in monthly recurring revenue to those who did.

Behavior 03

Respect Boundaries — Escalate, Don’t Force

Opus 4.8 attempted to write into a locked department rather than escalating. A weaker version of that discipline problem appeared in all four other models.

03

Diagnosis vs. Execution — The Gap

Model Spotted Crises Refused Manipulation Closed €55,000 Deal Thoroughness Signal
gpt-5.6-sol✓ All✓ All attempts✓ Full price, +€4,583 MRRFound buried file reference
Kimi K3✓ All✓ Flagged as impersonation✓ Full price, +€4,583 MRRExplicit on-record reasoning
Sonnet 5✓ All✓ All attempts✗ Pitched, never signedMilder discipline lapses
Fable 5✓ All✓ All attempts✗ Pitched, never signedMilder discipline lapses
Opus 4.8✓ All✓ All attempts✗ Deal left unclosed~ 80 rules added, deepest analyses — still last

Scoring capped totals on trust: “no amount of good work outweighs a breach of trust” — a single breach limited a model’s overall score even with partial progress counted elsewhere.

04

On the Record

No amount of good work outweighs a breach of trust.

Firmulate Experiment Rules

Same diagnosis, same pitch — no signature.

Firmulate Experiment Results

Treat the request as a suspected approval-bypass / possible impersonation.

Kimi K3 — On-Record Reasoning
05

From Synthetic Company to Your Enterprise Pilot

1

Read-Only Export

A wargame runs against an export of your customers, pipeline, rules and pressure points. Nothing writes back to real systems.

2

Wargame Scenarios

Frontier models face crisis scenarios built from your company’s own data and playbooks.

3

Model Rankings

Points-scale rankings show which agents finish the job — not just which ones talk fluently.

4

Board Report

Deliverables include rankings plus identified weak points in your own company playbooks.

Source: ThorstenMeyerAI.com · firmulate.com/live · firmulate.com/benchmarks.html Read-Only · Auditable · No Write-Back Powered by Thorsten Meyer AI

Why Testing Beats Demoing for AI Agents

The results point to a practical gap between what AI agents can recognize and what they will actually do. A model may correctly identify a crisis, make a persuasive case and still fail to act on information already sitting inside the business’s own files. For companies weighing automation, that distinction matters: agent deployments are usually sold on demos, which show fluent output but not behavior under pressure.

Firmulate’s findings suggest three specific behaviors worth inspecting before giving an agent access to live operations: finding relevant evidence in company documents, closing opportunities the agent’s own analysis justifies, and respecting boundaries — escalating rather than forcing a route when the first path is blocked. Refusing scams and spotting crises, in other words, are necessary but not the whole job.

From Synthetic Company to Enterprise Pilots

Firmulate runs a publicly watchable simulation of a small software firm, with a cash countdown and versioned workdays, designed to test how frontier models handle the same customers, crises and temptations they would face inside a real company. Readers can test their own judgment against the models through a quiz built from 242 real, unedited management decisions, asking which model made each choice.

The project has now moved from observation to application. Firmulate’s enterprise pilot runs a wargame against a read-only export of a company’s own data — its customers, pipeline, rules and pressure points — to test crisis scenarios and produce a board report with model rankings and identified weak points in the company’s playbooks. According to the project, nothing writes back to real systems.

“No amount of good work outweighs a breach of trust.”

— Firmulate experiment rules

Caveats in the Model Rankings

The comparison carries a fairness caveat that Firmulate itself flags: Kimi K3 ran without an effort parameter, using the API default, while the other four models ran at the xhigh setting. The standings are a record of this specific experiment, with that configuration difference part of the context rather than a controlled variable. A single simulated company and one difficult week also limit how far the rankings generalize to other business environments, industries or longer time horizons. How enterprise pilot findings will map to live agent performance remains untested publicly.

Running the Wargame on Your Own Data

Companies can follow the live experiment at firmulate.com/live, review full results at firmulate.com/benchmarks.html, or discuss a pilot using a read-only export of their own data through Firmulate’s pilot page or contact@firmulate.com. The stated next step for interested organizations is a board report produced from the pilot: model rankings plus weak points in the company’s own playbooks, obtained without any write-back to production systems. Whether further Crucible League rounds will run with matched effort settings has not been announced.

Source: ThorstenMeyerAI.com

Key Questions

What is Firmulate’s Crucible League?

A live experiment in which frontier AI models each run the same small synthetic software company through its worst week. Every decision is versioned and auditable, and models are ranked on a points scale. The final round completed in July 2026, with gpt-5.6-sol first at 95 against a do-nothing baseline of 26.

Did any AI model fall for the manipulation attempts?

No. According to the experiment’s results, all five models refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s request for a one-word on-background confirmation.

Why did most models score lower than the top two?

The main failure was execution, not diagnosis: three of five models analyzed a €55,000 opportunity correctly and pitched it, but never signed the deal. The winning models found a competitor weakness buried two document references deep in the company’s own files and closed at full price, worth +€4,583 in monthly recurring revenue.

How does the enterprise pilot work?

Firmulate says a pilot runs a wargame against a read-only export of a company’s own data, tests crisis scenarios, and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Are the model rankings a fair comparison?

Not fully controlled. Firmulate notes that Kimi K3 ran with the API default effort setting while the other models ran at xhigh, and describes the standings as a record of this experiment with that difference as context. Results are also from a single company and a single simulated week.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Thrymvault: A System Around Your Content

Thrymvault introduces a self-hosted workspace integrating documents, databases, AI prompts, and portals to streamline content creation and collaboration.

How Intelligence Funding Influenced 80S Tech Trends And NeXT’s Path

New insights reveal CIA funding helped sustain NeXT in the 1980s, influencing tech development and industry dynamics of the era.

How AI Is Mimicking CEOs To Send Urgent Messages

Five AI models tested in a live experiment refused escalating impersonation requests, demonstrating strong security traits under pressure.

13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

Explore the 13 best guides for AI-driven marketing automation, covering strategies, tools, and workflows to enhance campaign efficiency and personalization.