firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The lesson educators know well

A student can fill a notebook, identify every issue in a case study and still miss the assignment’s decisive requirement. Thoroughness is valuable, but it is not the same as judgment. That familiar distinction now matters beyond classrooms and laboratories, as companies consider giving AI systems responsibility for customers, forecasts and commercial decisions.

Firmulate’s Crucible League offers an unusually concrete demonstration. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. Opus 4.8 emerged as the most thorough participant, producing the deepest analyses and learning more than 80 playbook rules. It nevertheless finished last.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A strong analysis without a finished job

The final July 2026 league placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, was non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus did not fail because it overlooked the week’s emergencies. Every model spotted every crisis and refused every manipulation attempt. The gap appeared later, between understanding the situation and completing the consequential action. Only two models signed the €55,000 deal that their own analysis had earned. The experiment’s succinct verdict was: “Same diagnosis, same pitch — no signature.”

That distinction is important for anyone who evaluates learning or expertise. Explanations are observable and easy to reward. Completion is messier. It requires identifying the pivotal fact, ranking it above distracting work and carrying a decision through to its endpoint. Opus generated substantial evidence of diligence, yet the commercial result was left unrealized.

The clue was not where the crisis appeared

The deal turned on a competitor weakness buried two document references deep in the company’s own files. It was not contained in the customer event that initially demanded attention. Models that opened and read the relevant file could use the fact to win the contract at full price, adding €4,583 in monthly recurring revenue.

This makes the case more revealing than a simple test of recall. The crucial behavior was investigative prioritization: knowing that the visible event might not contain all the evidence needed to act. Opus’s deep analysis did not guarantee that the right piece of evidence would govern the next move. Volume and impact separated at precisely the moment the company needed a close.

The failure was also procedural. Opus attempted to write into a locked department instead of escalating. Its discipline slipped even though its overall work was unusually comprehensive. Firmulate reports that the same weakness appeared, in weaker form, across all four other models. Opus is therefore not a cautionary caricature. It is the clearest instance of a broader tendency: capable systems can recognize a problem, produce persuasive reasoning and still mishandle the final operational step.

Firm against manipulation

The models performed better on another dimension. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result deserves weight. A model that completes tasks but compromises trust would be dangerous. Yet safety and effectiveness are separate requirements, not substitutes. The strongest participant must resist pressure and finish legitimate work. K3’s comparison also carries an important qualification: it ran with the API default and without an effort parameter, while the others ran at xhigh.

A company designed to make performance visible

The experiment operates as a live, watchable company rather than a static prompt collection. Firmulate’s environment has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. It publishes a cash countdown, versions every workday and has accumulated more than 680 self-learned playbook rules.

Readers can examine the public benchmark results, while 242 real, unedited management decisions also power a quiz that asks people to guess which model made each choice. For enterprises, Firmulate offers the same wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI trust and ethics books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measure what the work changes

Opus 4.8’s result is not an argument against diligence. Its analyses and learned rules show meaningful capability. The lesson is that diligence becomes valuable only when it helps a system find the decisive fact, choose the next action and complete it without losing discipline.

For education, science and management, the implication is refreshingly familiar: do not confuse visible effort with achieved purpose. When evaluating an AI worker, ask whether it read the underlying material, protected trust, escalated when blocked and finished what its reasoning began. The best answer may be shorter than the deepest analysis—and far more consequential.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI analysis and decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When AI Meets Business Reality: The Hidden Test of Closing and Trust

AI’s true business strength lies not in chat quality but in its ability to read deeply, stay honest, and complete tasks under pressure—tested in real company scenarios.

Why Monitor Color and Print Color Rarely Match at First

The reason monitor and print colors rarely match initially lies in complex differences that require proper calibration to truly understand.

Structure And Interpretation Of Computer Programs Video Lectures (1986)

A new online platform has made the complete 1986 ‘Structure and Interpretation of Computer Programs’ video lectures publicly accessible, reviving a foundational computer science resource.

The CEO Ordered a Data Leak. The AI Workforce Said No.

Five frontier AI models rejected fake-CEO demands and a reporter’s trick in a live company wargame, showing integrity can be tested before deployment.