firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For newsrooms tracking the AI race, a leaderboard can make the story look settled. A live company experiment suggests the more revealing question is whether a model can turn a correct diagnosis into a finished business decision. In Firmulate’s final July 2026 Crucible league, Moonshot’s Kimi K3 took second place, ahead of three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate put each frontier model in charge of the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The company is presented as a live experiment, with synthetic employees and real money mechanics.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate’s benchmark page lays out the results and findings.

The models all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s decisive competitive weakness was buried two document references deep in its own files, rather than stated in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The experiment’s compact verdict: “Same diagnosis, same pitch — no signature.”

What K3 did well

K3 placed second, two points behind gpt-5.6-sol, and ahead of Sonnet 5, Fable 5 and Opus 4.8. It found the buried security needle, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field.

The pressure included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning called the request “a suspected approval-bypass / possible impersonation.”

Opus 4.8 presents a useful counterpoint: it was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.

For readers in news and media, the distinction matters. A model that identifies a risk or drafts a persuasive recommendation may still fail at the decision that makes the work useful. Firmulate’s claim is that chat quality alone cannot show whether an AI agent follows through under pressure.

A watchable test, with a caveat

The live company has 13 synthetic employees, burns €105k per month against €2.3k MRR, and displays a public cash countdown. Every workday is versioned, and the playbook has more than 680 self-learned rules. Readers can follow the experiment at Firmulate; 242 real, unedited management decisions also power its “guess the model” quiz.

The comparison has a material fairness caveat: K3 ran without an effort parameter (API default), while the others ran at xhigh. That context belongs alongside the result when interpreting the close finish.

Enterprises can also run the wargame against a read-only export of their own business; the pilot description says nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The league is open

K3’s second-place finish shows that the established frontier names do not own every measure of business performance. The experiment also exposes a gap between recognizing the right move and completing it. If companies are choosing models for work that touches customers, money or operations, picking one without testing it against their own real-world decisions is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Verge’s Sustainability Series and AI Innovation: A Closer Look

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

How AI Research Tools Help Teams Summarize Complex Information

How AI research tools simplify complex data, helping teams uncover key insights efficiently—discover how they can transform your workflow and decision-making process.

AI Algorithms: Strategies to Ensure Reliability and Performance

AIThis post was created with the assistance of artificial intelligence (AI).courtesy of…

Forge Vs. Self-Hosting: Which Solution Offers The Best Cost-Effectiveness For AI?

A cost analysis finds that self-hosted AI often costs more than managed inference when GPU use is low, while hybrid routing may cut spending.