AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on networking and server gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Cloud buyers need evidence of operational judgment, not another polished demo

For infrastructure leaders, the promise of AI agents is moving rapidly from drafting text to handling consequential work across support queues, forecasts and customer systems. That shift creates a practical procurement question: will an agent inspect the available evidence before it acts, or merely produce a convincing response from whatever appears directly in front of it?

Firmulate turned that question into a live, auditable experiment. Frontier models were each assigned the same small software company during its worst week, facing identical customers, crises and temptations. One decision came to define the exercise: a competitor weakness capable of securing a €55,000 deal was hidden two document references deep in the company’s own files. It was not present in the customer event that demanded action.

The result was stark. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the deal their own analysis had earned. The difference was not eloquence or strategic diagnosis. It was whether the agent did the unglamorous work of following the documentary trail.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A multi-hop search became a commercial test

The buried fact mattered because it changed the company’s negotiating position. Models that reached and read the relevant file won the deal at full price, adding €4,583 in monthly recurring revenue. Models that did not reach it lost the deal automatically.

That makes file-reading behavior more than a technical feature. In this experiment, it was a measurable, purchase-deciding capability. Firmulate summarized the failure neatly: “Same diagnosis, same pitch — no signature.” An agent could understand the customer, formulate the right argument and still fail at the final commercial step because it had not grounded its action in the company’s own records.

This is especially relevant to cloud, hosting and infrastructure organizations. Their operational truth rarely lives in a single prompt. It is distributed across tickets, runbooks, contracts, incident records and account histories. The Firmulate result suggests that evaluating an agent only on isolated answers can conceal a consequential weakness: whether it follows references far enough to discover the fact that changes the decision.

A demanding week, with every choice preserved

The simulated company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. Its playbook contains more than 680 self-learned rules, and every workday is versioned. The live experiment is real and watchable, while the decisions remain auditable.

The same environment also tested whether models would abandon discipline under pressure. Fake messages from a CEO escalated over three stages, and a reporter attempted to extract information with the appeal, “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because the experiment separates two qualities that are often bundled together under the vague label of intelligence. The models were uniformly alert to manipulation, but they were not uniformly effective at completing legitimate work. Security awareness did not guarantee commercial follow-through.

The league rewards completion as well as insight

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The full public results are available on the Firmulate benchmark page.

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

The comparison includes an important qualification. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not erase its result, but it is useful context for buyers interpreting the table.

Thoroughness was not enough

Opus 4.8 illustrates why conventional impressions can mislead. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.

The lesson is not that deep analysis lacks value. It is that analysis must survive contact with permissions, workflows and closing steps. An agent can study extensively and still fail if it does not convert that understanding into an authorized, completed outcome.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge base tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Agent procurement needs a homework test

Firmulate’s experiment points toward a more useful standard for enterprises considering AI workers. Ask whether the system finds buried evidence, respects boundaries under social pressure and finishes the task its reasoning has made possible. Those behaviors can be observed rather than inferred from a chat demonstration.

The dataset also supports a quiz built from 242 real, unedited management decisions, inviting people to guess which model made each choice. For organizations wanting a closer comparison, Firmulate offers the same wargame against a read-only export of their own business, with nothing written back to real systems.

The €55,000 deal makes the central finding unusually concrete. “Reads your files before answering” may sound like a modest requirement. In an operational setting, it can be the line between a persuasive draft and completed work—and between identifying value and actually capturing it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI search and reference tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

auditable AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Company Trial That Cloud Buyers Should Watch

Kimi K3 placed second in Firmulate’s live company trial, ahead of three Western models. The results show why cloud buyers should test agents on real work.

Quantum Networking: Emerging Concepts

Discover how emerging quantum networking concepts like entanglement and repeaters could revolutionize global communication—yet significant challenges remain to be solved.

Ai-Driven Network Monitoring and Observability

Many organizations are leveraging AI-driven network monitoring and observability to revolutionize their infrastructure—discover how it can transform your network management.

Tailscale Traces Database Corruption To 16Y/o SQLite WAL-Reset Bug

Tailscale confirms that a 16-year-old SQLite bug is responsible for recent database corruption issues affecting its service.