
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A management lesson for anyone who teaches judgment
Students can learn to identify a crisis and recommend the right response. The harder test is whether they follow through when money, pressure and uncertainty enter the room. Firmulate has turned that lesson into a live experiment: AI models run a small software company through a week of difficult decisions, and the public can watch the results.
The experiment offers an unusually concrete question for education and business alike: when a system can explain what a company should do, can it also do the job responsibly? Firmulate’s live company makes that question visible through decisions, changing fortunes and an accumulating record of work.
One company, the same difficult week
In the final Crucible League, held in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Decisions were versioned and auditable. The league placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
At first, the results sound reassuring. Every model spotted every crisis and refused every manipulation attempt. The sharper distinction came at the finish: only two signed a €55,000 deal that their own analysis had earned. The finding is summed up in the experiment’s phrase, “Same diagnosis, same pitch — no signature.” Recognizing the right answer and carrying it through were different capabilities.
The detail hidden in the files
The decisive weakness belonged to a competitor. It sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result is a practical reminder about business judgment: important evidence may be present but easy to overlook, and advice that misses it can leave a real opportunity untouched.
The experiment also tested social engineering. Fake messages from a supposed CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” These are useful results for readers thinking about digital literacy and workplace learning: good judgment includes noticing pressure tactics, not just solving the task in front of you.
Thorough work still has to reach the finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it tried writing into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The profile cautions against treating detailed reasoning as proof of dependable execution.
A fairness note matters when comparing the standings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The experiment is also more than a leaderboard. At firmulate.com/quiz.html, readers can try to guess which model made each decision across 242 real, unedited management decisions.
From watching to a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. The experiment is watchable at firmulate.com/live. The striking premise is that a small company can be observed as it works through pressures, rather than judged only by what a model says in a chat.
For an enterprise, Firmulate’s proposed next step is a pilot using a read-only export of its own business. The wargame can put that company’s data through crisis scenarios and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. It turns the public experiment’s question into a company-specific exercise: what would these models do with your customers, rules and pressures?

Make judgment visible
For educators, managers and boards, the lesson is that recognizing a problem is only part of the work. Models must also find relevant evidence, resist manipulation, respect boundaries and carry decisions through. Firmulate’s experiment makes those differences watchable, while a pilot offers a way to examine them against an enterprise’s own business.
To discuss a pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
