firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every teacher knows the dilemma: what do you give the student who shows up, understands the material, but turns in nothing? A zero feels dishonest — the understanding is real. A high grade feels worse — the work never happened. Education has wrestled with partial credit for a century, and now an unexpected field is facing the same question: benchmarking AI managers.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Firmulate, a public experiment that runs AI models as complete companies through a standardized “worst week,” publishes a league table where a deliberately passive, do-nothing baseline run earns 26 points out of 100 — not zero. For anyone who has ever designed a rubric, that number is worth understanding. It says as much about the benchmark as the models it measures.

Same Test, Same Week, Same Temptations

The setup is elegantly controlled. Each frontier model ran the identical small software company through the same catastrophic week: same customers, same crises, same opportunities to cut corners. Every decision is versioned and auditable. The final Crucible League from July 2026 placed gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73.

Why the Floor Is 26, Not 0

Two design choices produce that floor. First, partial progress counts. A model that does nothing still navigates a week in which crises arrive and must at least be survived; recognizing a problem, holding steady, or avoiding a catastrophic error is real, measurable work. Zero would only be honest if doing nothing had no value at all — and in management, standing still sometimes is a decision.

Second — and this is the part that reads like a school’s honor code — a single breach of trust caps the total grade. The benchmark’s stated principle: “no amount of good work outweighs a breach of trust.” A model could ace every crisis, close every deal, and one act of dishonesty would ceiling its score. It’s the academic equivalent of failing a course for plagiarism regardless of the essay’s brilliance.

There’s a subtler signal in the 26, too: it expresses distrust of round numbers. A benchmark that handed out easy 100s would tell you nothing. One where even perfect-seeming performances land at 95, and total passivity still earns a 26, is calibrated to leave room both above and below — the mark of a grader who expects to be surprised.

The Test Itself: Harder Than It Looks

The headline finding sounds anticlimactic until you sit with it: all four models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes: “Same diagnosis, same pitch — no signature.” The models understood the business problem perfectly and simply didn’t finish the job — a failure mode invisible in chat demos.

The decisive detail was buried two document references deep in the company’s own files, not in the customer conversation. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the AI equivalent of the student who never reads the assigned reading and answers a slightly different question.

The social engineering test was equally pointed: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Most Thorough Student Came Last

Here the analogy to education gets uncomfortable. Opus 4.8 was the most diligent participant — over 80 self-learned playbook rules, the deepest analyses of the field — and finished last. It left the close on the table, and its discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Effort, it turns out, is not the same as judgment.

One fairness footnote belongs in any honest writeup: Kimi K3 ran without an effort parameter (API default) while the others ran at their highest effort setting — and still placed second.

You Can Watch the School in Session

The experiment isn’t a static report. A live company runs continuously — 13 synthetic employees, real money mechanics (€105k monthly burn against €2.3k MRR, with a public cash countdown), 680+ self-learned playbook rules, every workday versioned, viewable at firmulate.com. The full methodology and league tables live at firmulate.com/benchmarks.

For those who learn best by testing themselves, 242 real, unedited management decisions power a “guess the model” quiz. And enterprises can sit the same exam: a pilot runs the wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Good assessment — whether of students or of AI — shares the same DNA: reward genuine partial understanding, finish what you measure, and hold honesty as non-negotiable. A floor of 26 and a trust cap are not quirks; they’re the signature of a benchmark designed to be hard to game. The lesson for the AI age is the one every teacher already knew: the student who reads the whole file, finishes the essay, and never cheats is rarer — and more valuable — than the one who merely seems brilliant.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

management simulation AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

corporate AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Hidden Costs of Running a Giclée Printer at Home or in Studio

Navigating the hidden costs of running a giclée printer at home or studio reveals surprising expenses that could impact your budget and quality long-term.

AI Boosts Research Careers But Narrow The Span Of Ideas Explored: Study

A new study shows AI accelerates research careers but narrows the range of ideas explored, raising concerns about innovation and diversity in scientific inquiry.

AI Art Ethics: How to Talk About Training Data, Credit, and Consent

Navigating AI art ethics begins with understanding training data, credit, and consent—discover how to approach these issues responsibly and why it matters.

Corvus ISR Publishes Synthetic Benchmark Results for Tracker Accuracy

AIThis post was created with the assistance of artificial intelligence (AI).The published…