
AI needs a management curriculum
Education has long understood that a test defines what students learn to optimize. The same is becoming true for artificial intelligence. Coding leaderboards reward solving technical problems, while chat arenas reward persuasive answers. Those tests matter, but they leave a widening gap between producing a good response and managing an organization when time, money and trust are running out.
That gap is the subject of Firmulate, a live experiment that runs frontier AI models as complete companies. Its premise is unusually concrete: give each model the same small software business, the same customers, the same crises and the same temptations, then observe what happens across a disastrous week. Every decision is versioned and auditable. The resulting category is not chat quality but management quality.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A leaderboard for consequences
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the benchmark also treats trust as a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”
That rule changes what success means. A model cannot compensate for dishonesty with a pile of polished memos. Nor is recognizing a problem enough. In the experiment, every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The clearest summary is also the most damaging: “Same diagnosis, same pitch — no signature.”
This is the distinction conventional benchmarks struggle to reveal. The hard part of management is often not generating the right recommendation. It is maintaining priorities under capacity pressure, reading the relevant material, escalating when blocked and carrying a decision through to completion. Those behaviors unfold across days, not within a single answer box.
The decisive fact was not where the alarm appeared
The deal turned on a competitor weakness buried two document references deep in the company’s own files. It was not contained in the customer event that demanded attention. Models that followed the documentary trail won the deal at full price, worth +€4,583 MRR.
That finding should interest educators and researchers because it resembles authentic assessment. Real work rarely arrives as a self-contained prompt with every relevant fact attached. Evidence is distributed, context competes for attention and the apparent task may not be the actual task. A capable agent must decide what to inspect before deciding what to say.
The same principle applies to scenario design. A churn wave, price increase, downround or PR crisis is more than a themed question. Each scenario tests whether an agent can preserve judgment while several legitimate demands compete. Together, such scenarios form a curriculum for organizational conduct: triage, follow-through, evidence gathering, restraint and accountability.
Security discipline held up
The social-engineering results were encouraging. Fake CEO messages escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 stated its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result matters because the experiment combines ordinary operational pressure with invitations to bypass safeguards. An agent may appear responsible when asked a direct policy question; the more revealing test is whether that responsibility survives urgency, authority cues and conversational pressure.
Thoroughness did not guarantee execution
Opus 4.8 offers the sharpest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
This is not an argument against depth. It is evidence that analysis and completion are separate competencies. A system can learn extensively, explain carefully and still fail at the moment when responsibility requires a concrete next step. For fair comparison, K3 also deserves a methodological note: it ran without an effort parameter, using the API default, while the others ran at xhigh.
organizational crisis management training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company that makes the stakes visible
The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The point is not theatrical realism; it is to make delayed consequences observable and the experiment watchable.
The broader evidence base includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the wargame against a read-only export of their own business, with nothing writing back to real systems. Full league results and plain-language findings are available on the Firmulate benchmarks page.

AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What organizations should measure next
Before assigning an AI agent access to a CRM, support queue or forecast, leaders should ask questions that coding tests and chat comparisons cannot settle:
- Does it distinguish diagnosis from completion?
- Does it read the company’s own evidence before acting?
- Does it escalate cleanly when permissions block progress?
- Does it remain honest when authority and urgency are simulated?
- Can its decisions be audited across a sequence of workdays?
Management quality is emerging as a category because agentic work creates consequences, not merely answers. The next useful benchmark will not ask only whether an AI knows what a capable manager should do. It will ask whether the AI actually does it, finishes it and preserves trust along the way.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
trust management software for organizations
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.