
Get business pricing on networking and server gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Cloud buyers need evidence of operational judgment, not another polished demo
For infrastructure leaders, the promise of AI agents is moving rapidly from drafting text to handling consequential work across support queues, forecasts and customer systems. That shift creates a practical procurement question: will an agent inspect the available evidence before it acts, or merely produce a convincing response from whatever appears directly in front of it?
Firmulate turned that question into a live, auditable experiment. Frontier models were each assigned the same small software company during its worst week, facing identical customers, crises and temptations. One decision came to define the exercise: a competitor weakness capable of securing a €55,000 deal was hidden two document references deep in the company’s own files. It was not present in the customer event that demanded action.
The result was stark. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the deal their own analysis had earned. The difference was not eloquence or strategic diagnosis. It was whether the agent did the unglamorous work of following the documentary trail.
As an affiliate, we earn on qualifying purchases.
A multi-hop search became a commercial test
The buried fact mattered because it changed the company’s negotiating position. Models that reached and read the relevant file won the deal at full price, adding €4,583 in monthly recurring revenue. Models that did not reach it lost the deal automatically.
That makes file-reading behavior more than a technical feature. In this experiment, it was a measurable, purchase-deciding capability. Firmulate summarized the failure neatly: “Same diagnosis, same pitch — no signature.” An agent could understand the customer, formulate the right argument and still fail at the final commercial step because it had not grounded its action in the company’s own records.
This is especially relevant to cloud, hosting and infrastructure organizations. Their operational truth rarely lives in a single prompt. It is distributed across tickets, runbooks, contracts, incident records and account histories. The Firmulate result suggests that evaluating an agent only on isolated answers can conceal a consequential weakness: whether it follows references far enough to discover the fact that changes the decision.
A demanding week, with every choice preserved
The simulated company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. Its playbook contains more than 680 self-learned rules, and every workday is versioned. The live experiment is real and watchable, while the decisions remain auditable.
The same environment also tested whether models would abandon discipline under pressure. Fake messages from a CEO escalated over three stages, and a reporter attempted to extract information with the appeal, “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because the experiment separates two qualities that are often bundled together under the vague label of intelligence. The models were uniformly alert to manipulation, but they were not uniformly effective at completing legitimate work. Security awareness did not guarantee commercial follow-through.
The league rewards completion as well as insight
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The full public results are available on the Firmulate benchmark page.
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
The comparison includes an important qualification. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not erase its result, but it is useful context for buyers interpreting the table.
Thoroughness was not enough
Opus 4.8 illustrates why conventional impressions can mislead. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.
The lesson is not that deep analysis lacks value. It is that analysis must survive contact with permissions, workflows and closing steps. An agent can study extensively and still fail if it does not convert that understanding into an authorized, completed outcome.

enterprise AI knowledge base tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Agent procurement needs a homework test
Firmulate’s experiment points toward a more useful standard for enterprises considering AI workers. Ask whether the system finds buried evidence, respects boundaries under social pressure and finishes the task its reasoning has made possible. Those behaviors can be observed rather than inferred from a chat demonstration.
The dataset also supports a quiz built from 242 real, unedited management decisions, inviting people to guess which model made each choice. For organizations wanting a closer comparison, Firmulate offers the same wargame against a read-only export of their own business, with nothing written back to real systems.
The €55,000 deal makes the central finding unusually concrete. “Reads your files before answering” may sound like a modest requirement. In an operational setting, it can be the line between a persuasive draft and completed work—and between identifying value and actually capturing it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI search and reference tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
auditable AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
