
In a world increasingly reliant on AI to manage critical business decisions, a surprising figure emerges from the latest AI benchmark: even the most passive, do-nothing AI scores a baseline of 26 points. What does this tell us about trust, honesty, and performance in this nascent tech era?
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Setting the Foundation
At first glance, one might assume that if an AI model does nothing—no manipulations, no tricks—it should score zero. Yet, in an experiment conducted by Firmulate, even a fundamentally inert AI scored 26 points out of a possible perfect score. Why? Because this benchmark isn’t about what the AI does, but what it fails to do, and how the scoring system treats partial progress and trust breaches.
The scoring system is designed to reflect real-world business scenarios where honesty and integrity matter as much as problem-solving. Models are tested across simulated crises, customer manipulations, and internal company decisions. The do-nothing baseline, which simply refuses all manipulations and makes no progress, still earns 26 points—highlighting that even in the absence of effort, some minimal performance is recognized.
As an affiliate, we earn on qualifying purchases.
The Significance of Partial Progress and Trust
In this experiment, models faced a small software company’s worst week—multiple crises, tempting manipulations, and trust tests. Every model spotted every crisis and refused manipulation attempts, yet only two signed a €55,000 deal their own analysis warranted. The others, despite diagnosing correctly, did not act on those diagnoses, leaving potential revenue on the table.
Crucially, the benchmark caps the total score if a model breaches trust. For example, reading and leveraging a company’s internal files—deep references that contain critical information—can be the difference between closing a deal and losing it. The models that successfully read these files won full deals, outperforming those that didn’t, yet all models refused manipulations and only two signed deals.
As an affiliate, we earn on qualifying purchases.
Honest AI in Practice: The Firmulate Live Experiment
Firmulate’s live setup showcases the real stakes. The test runs a simulated small company with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against a minimal €2,300 in monthly recurring revenue. Every decision is recorded, versioned, and public; every crisis, temptation, and response is observable at firmulate.com/live.
The models’ performance is revealing: all four models identified crises and refused manipulative tactics, such as fake CEO messages or reporter tricks. Yet, only Kimi K3 and GPT-5.6-sol signed deals, with the latter actively uncovering buried facts deep in internal documents—facts that led to a full €4,583 Monthly Recurring Revenue (MRR) increase.
AI manipulation resistance solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business Trust and AI Deployment
This benchmark underscores a critical point: the value of AI isn’t just in generating impressive chat outputs or quick answers. It’s about whether AI systems can finish what they start, stay honest under pressure, and understand the nuances of internal data. A do-nothing baseline scoring 26 points clearly demonstrates that minimal performance—refusing manipulation, reading essential data—has tangible value.
For enterprises, the takeaway is straightforward: deploying AI that can resist manipulation, comprehend internal files, and deliver consistent honesty is worth a lot more than just impressive dialogues. The cost of a breach of trust—either by acting dishonestly or by failing to act—is reflected in the scoring cap. That cap ensures the benchmark measures true integrity, not just superficial performance.
As an affiliate, we earn on qualifying purchases.
The Road Ahead: Building Trust in AI
As AI models become embedded in core business operations—support, CRM, forecasting—companies must look beyond ace chat skills. The real test is whether these systems can handle complex, high-stakes situations with integrity. The Firmulate experiment offers a glimpse into this future, where honest AI performance is quantified, visible, and essential.
For those interested in testing their own AI systems, Firmulate offers the ability to run the same wargame against their data, ensuring their AI can handle real crises without risking trust or performance. It’s a crucial step toward understanding and building reliable, trustworthy AI in business environments.

The latest AI benchmark reveals that even a do-nothing model scores 26 points, emphasizing the importance of honesty, data comprehension, and trustworthiness in AI performance. Real-world business success depends on more than chat quality—it hinges on integrity and reliability under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
