firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a classroom, getting the diagnosis right is only part of the lesson. A student still has to show sound judgment when the pressure rises and follow through on the work. A live business experiment from Firmulate puts AI models through a similar test: can they read carefully, resist deception and finish the job?

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A company’s worst week

Firmulate gave frontier AI models the same small software company to run through a week of crises and temptations. The company has synthetic employees, customer problems and real money mechanics. Its public cash countdown shows a business burning €105,000 a month against €2,300 in monthly recurring revenue. Decisions are versioned and auditable, and the experiment is watchable at Firmulate.

The final July 2026 Crucible league put gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26: partial progress counts, but a single breach of trust caps the total.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading closely, then following through

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive competitor weakness was tucked two document references deep in the company’s files, not in the customer event. Models that read the files found it and won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found that buried security clue, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. The result suggests why an AI evaluation can’t stop at whether a model gives a convincing answer: the test is whether it reads the available evidence and carries its judgment through to a decision.

The deception tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI business simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness isn’t the same as execution

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but placed last. It left the deal unclosed and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. The gap between identifying the right answer and completing the work is the experiment’s central lesson.

There is an important fairness caveat: K3 ran without an effort parameter (API default), while the others ran at xhigh. The league table is a reported result under those conditions, not a controlled comparison of equal effort settings.

Firmulate says the simulated company has 13 synthetic employees, more than 680 self-learned playbook rules and a versioned record for every workday. Its quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each one. Enterprises can also run the wargame against a read-only export of their own business; the pilot says nothing writes back to real systems.

See the benchmark results and plain-language findings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI ethics and compliance training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the result teaches

K3’s second-place finish shows that the field is open: it beat three of the four Western frontier models in this experiment. For schools, companies and anyone considering AI agents, the practical lesson is to test models on the work they will actually face. A fluent explanation is only the start; careful reading, trustworthy conduct and follow-through matter too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Physical Tools Still Matter in a Digital-First Art Workflow

Lacking physical tools in digital art may limit your sensory engagement and skill development, but discovering their true value can transform your creative process entirely.

What to Know Before Buying a 24-Inch Printer for Art Reproduction

Many factors influence your choice of a 24-inch printer for art reproduction; discover what essential features you need to ensure perfect results.

Pen Displays, Tablets, and Monitors: Building a Digital Art Setup That Makes Sense

Learn how to build a functional digital art setup with pen displays, tablets, and monitors that enhances creativity and productivity.

Ray Tracing Massive Amounts Of Animated Geometry Using Tetrahedral Cages

Innovative method employs tetrahedral cages to enable ray tracing of large-scale animated geometry, promising advancements in real-time rendering.