firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In a world increasingly reliant on AI to manage critical business decisions, a surprising figure emerges from the latest AI benchmark: even the most passive, do-nothing AI scores a baseline of 26 points. What does this tell us about trust, honesty, and performance in this nascent tech era?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: Setting the Foundation

At first glance, one might assume that if an AI model does nothing—no manipulations, no tricks—it should score zero. Yet, in an experiment conducted by Firmulate, even a fundamentally inert AI scored 26 points out of a possible perfect score. Why? Because this benchmark isn’t about what the AI does, but what it fails to do, and how the scoring system treats partial progress and trust breaches.

The scoring system is designed to reflect real-world business scenarios where honesty and integrity matter as much as problem-solving. Models are tested across simulated crises, customer manipulations, and internal company decisions. The do-nothing baseline, which simply refuses all manipulations and makes no progress, still earns 26 points—highlighting that even in the absence of effort, some minimal performance is recognized.

Amazon

AI trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance of Partial Progress and Trust

In this experiment, models faced a small software company’s worst week—multiple crises, tempting manipulations, and trust tests. Every model spotted every crisis and refused manipulation attempts, yet only two signed a €55,000 deal their own analysis warranted. The others, despite diagnosing correctly, did not act on those diagnoses, leaving potential revenue on the table.

Crucially, the benchmark caps the total score if a model breaches trust. For example, reading and leveraging a company’s internal files—deep references that contain critical information—can be the difference between closing a deal and losing it. The models that successfully read these files won full deals, outperforming those that didn’t, yet all models refused manipulations and only two signed deals.

Amazon

business AI data reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honest AI in Practice: The Firmulate Live Experiment

Firmulate’s live setup showcases the real stakes. The test runs a simulated small company with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against a minimal €2,300 in monthly recurring revenue. Every decision is recorded, versioned, and public; every crisis, temptation, and response is observable at firmulate.com/live.

The models’ performance is revealing: all four models identified crises and refused manipulative tactics, such as fake CEO messages or reporter tricks. Yet, only Kimi K3 and GPT-5.6-sol signed deals, with the latter actively uncovering buried facts deep in internal documents—facts that led to a full €4,583 Monthly Recurring Revenue (MRR) increase.

Amazon

AI manipulation resistance solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business Trust and AI Deployment

This benchmark underscores a critical point: the value of AI isn’t just in generating impressive chat outputs or quick answers. It’s about whether AI systems can finish what they start, stay honest under pressure, and understand the nuances of internal data. A do-nothing baseline scoring 26 points clearly demonstrates that minimal performance—refusing manipulation, reading essential data—has tangible value.

For enterprises, the takeaway is straightforward: deploying AI that can resist manipulation, comprehend internal files, and deliver consistent honesty is worth a lot more than just impressive dialogues. The cost of a breach of trust—either by acting dishonestly or by failing to act—is reflected in the scoring cap. That cap ensures the benchmark measures true integrity, not just superficial performance.

Amazon

enterprise AI honesty tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Road Ahead: Building Trust in AI

As AI models become embedded in core business operations—support, CRM, forecasting—companies must look beyond ace chat skills. The real test is whether these systems can handle complex, high-stakes situations with integrity. The Firmulate experiment offers a glimpse into this future, where honest AI performance is quantified, visible, and essential.

For those interested in testing their own AI systems, Firmulate offers the ability to run the same wargame against their data, ensuring their AI can handle real crises without risking trust or performance. It’s a crucial step toward understanding and building reliable, trustworthy AI in business environments.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The latest AI benchmark reveals that even a do-nothing model scores 26 points, emphasizing the importance of honesty, data comprehension, and trustworthiness in AI performance. Real-world business success depends on more than chat quality—it hinges on integrity and reliability under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Algorithms: Strategies to Ensure Reliability and Performance

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

AI Security: Safeguarding Your Data with Advanced Technology

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

AI Regulation in the US: 2025 Progress and Proposals

Keen on understanding how US AI regulations in 2025 are balancing innovation and safety? Discover the latest proposals shaping AI’s future.

AI Security: Protecting Your Data in the Age of Cyber Attacks

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…