firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For newsrooms, the question about AI is moving beyond whether a model can summarize a crisis. Can it notice what matters, resist pressure, and make the consequential call? Firmulate’s live experiment puts models through a company’s worst week, where a correct diagnosis does not always lead to action.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate runs AI models as complete companies, measuring management decisions rather than chat quality. In its final Crucible League, published in July 2026, each frontier model faced the same small software company, customers, crises, and temptations. Every decision was versioned and auditable.

The live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. The experiment is watchable at Firmulate.

Recognition is not execution

All the models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”

The buried detail helps explain the gap. A decisive weakness in a competitor was two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

The standings show that thoroughness alone did not secure the outcome. GPT-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Pressure, discipline, and a fairness caveat

Models faced fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it placed last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

There is a qualification to the league comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to testing your own business

For media organizations and other enterprises, the experiment points to a practical next step: test an AI workforce against your own company’s pressures before relying on it. Firmulate says a pilot can use a read-only export to create a digital twin, run crisis scenarios, and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

See the Firmulate pilot and contact contact@firmulate.com to discuss running the wargame against your business.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The takeaway

Firmulate’s experiment suggests that spotting the right answer is only part of management. Under pressure, models also have to find the evidence, follow through, and preserve trust. Enterprises can test those decisions against a read-only export of their own business, with no changes written to live systems. Explore the pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Keeps Your Data Safe While You Sleep? Unveiling AI Security

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

Microsoft Launches New Tool for Content Authentication and Election Support

AIThis post was created with the assistance of artificial intelligence (AI). Prime…

Mastering the Art of Privacy in AI: Essential Principles and Techniques

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

The Risks and Future of AI Security Unveiled

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…