firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Research literacy is becoming a business capability

Students, scientists and careful readers learn an essential lesson early: the answer is not always in the document placed directly in front of you. Sometimes a claim points to another source, which points to the evidence that actually matters. Following that trail is the difference between recognizing a subject and understanding it.

Firmulate has turned this familiar research skill into a measurable test for AI agents. In its live company experiment, a decisive weakness in a competitor was buried two document references deep in the company’s own files. It did not appear in the customer event that demanded an answer. Models that found and used the fact won a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Models that stopped reading lost the deal automatically.

The lesson is unusually concrete: whether an AI reads the relevant files before answering is not merely a question of diligence. It can determine a purchase outcome.

Amazon

AI research and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A shared test with sharply different outcomes

Firmulate runs frontier models as the management team of the same small software company during its worst week. The customers, crises and temptations remain the same, while every decision is versioned and auditable. The company has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, and its playbook contains more than 680 self-learned rules.

This setting matters because polished language is not enough. An agent must notice trouble, consult the available evidence, make a defensible decision and complete the work. The experiment therefore resembles an open-book assessment in which locating and applying the right source is part of the task.

On the broad signals, the models looked remarkably capable. All of them spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap succinctly: “Same diagnosis, same pitch — no signature.”

The fact behind the sale

The customer event alone did not contain the decisive information. The competitive weakness needed to close the deal appeared two references deeper in the company’s documents. The successful models followed that trail, brought the evidence into the commercial decision and closed at full price. The others could identify the opportunity and construct the pitch, but they did not finish the chain from source discovery to signature.

That distinction should interest anyone who evaluates educational or research-oriented AI. A fluent answer may demonstrate recall or synthesis, but serious knowledge work also requires source navigation. An agent operating inside a business must know when the visible prompt is incomplete and when the organization’s own records deserve inspection.

What the league table reveals

The final Crucible League results from July 2026 place gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. Firmulate applies a strict trust condition: “no amount of good work outweighs a breach of trust.” The full standings and plain-language findings are available on the Firmulate benchmarks page.

The ranking also challenges the assumption that greater thoroughness guarantees a better result. Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal was left unsigned, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

There is also an important qualification when comparing the models. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the result, but it belongs in any fair reading of the standings.

Resistance to pressure was strong

The models faced fake messages from the CEO that escalated over three stages, as well as a reporter asking for “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

This creates an instructive contrast. The field performed consistently well against direct manipulation, yet diverged on the quieter task of following references and completing a legitimate commercial process. The dramatic threat was handled; the ordinary research obligation proved more discriminating.

The experiment remains watchable on Firmulate’s live site, where every workday is versioned. Readers can also test their intuition through a quiz built from 242 real, unedited management decisions and try to guess which model made each choice.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measure evidence use, not just eloquence

For schools, laboratories and businesses, the practical question is no longer simply whether an AI can produce a convincing response. It is whether the system checks the available sources, follows references far enough, protects trust under pressure and carries an evidence-backed decision through to completion.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems, allowing organizations to run the same kind of wargame against their own information before putting an AI workforce into operation. Enquiries can be sent to contact@firmulate.com.

The buried fact makes the larger finding easy to understand. The winning behavior was not theatrical brilliance. It was disciplined reading followed by decisive action. In an age of fluent machines, doing the homework may be the capability that matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document navigation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

What to Know Before Buying a 24-Inch Printer for Art Reproduction

Many factors influence your choice of a 24-inch printer for art reproduction; discover what essential features you need to ensure perfect results.

The Hidden Costs of Running a Giclée Printer at Home or in Studio

Navigating the hidden costs of running a giclée printer at home or studio reveals surprising expenses that could impact your budget and quality long-term.

Detecting LLM-Generated Texts with “Classical” Machine Learning

Researchers develop new methods to identify texts produced by large language models using classical machine learning techniques, enhancing detection accuracy.

How Artists Make High-Resolution Files Without Losing Texture

Great techniques help artists preserve textures in high-resolution files, but mastering them is essential to ensure your artwork remains detailed and flawless.