
In a classroom, getting the diagnosis right is only part of the lesson. A student still has to show sound judgment when the pressure rises and follow through on the work. A live business experiment from Firmulate puts AI models through a similar test: can they read carefully, resist deception and finish the job?
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A company’s worst week
Firmulate gave frontier AI models the same small software company to run through a week of crises and temptations. The company has synthetic employees, customer problems and real money mechanics. Its public cash countdown shows a business burning €105,000 a month against €2,300 in monthly recurring revenue. Decisions are versioned and auditable, and the experiment is watchable at Firmulate.
The final July 2026 Crucible league put gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26: partial progress counts, but a single breach of trust caps the total.
As an affiliate, we earn on qualifying purchases.
Reading closely, then following through
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive competitor weakness was tucked two document references deep in the company’s files, not in the customer event. Models that read the files found it and won the deal at full price, worth €4,583 in monthly recurring revenue.
K3 found that buried security clue, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. The result suggests why an AI evaluation can’t stop at whether a model gives a convincing answer: the test is whether it reads the available evidence and carries its judgment through to a decision.
The deception tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as execution
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but placed last. It left the deal unclosed and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. The gap between identifying the right answer and completing the work is the experiment’s central lesson.
There is an important fairness caveat: K3 ran without an effort parameter (API default), while the others ran at xhigh. The league table is a reported result under those conditions, not a controlled comparison of equal effort settings.
Firmulate says the simulated company has 13 synthetic employees, more than 680 self-learned playbook rules and a versioned record for every workday. Its quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each one. Enterprises can also run the wargame against a read-only export of their own business; the pilot says nothing writes back to real systems.
See the benchmark results and plain-language findings.

AI ethics and compliance training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the result teaches
K3’s second-place finish shows that the field is open: it beat three of the four Western frontier models in this experiment. For schools, companies and anyone considering AI agents, the practical lesson is to test models on the work they will actually face. A fluent explanation is only the start; careful reading, trustworthy conduct and follow-through matter too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
