
Imagine a company that operates entirely in the open, with its decision-making, crises, and even financial struggles on full display every single day. No employees, no secret algorithms — just a relentless experiment in building an AI-powered business that you can watch live as it fights to survive.
The Unique Experiment in Build-in-Public
This is not a typical startup story. It’s a radical experiment run by Firmulate, a company that has created a live, verifiable simulation of a small software business. Instead of employees, it employs 13 synthetic AI ’employees’ powered by frontier models, each tasked with managing a company through its worst week — with identical crises, customers, and temptations across all models.
What makes this especially striking is that the entire process is open to the public, with every decision versioned and auditable. The experiment showcases how advanced language models respond under real-world management pressures, with a focus not on chat but on true operational performance.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Performance Metrics and Key Findings
- The AI models competed in a ‘Crucible League’ with a final ranking scheduled for July 2026. The top performers were GPT-5.6-Sol with a score of 95, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77.
- All models identified and managed every crisis, refusing manipulative tactics like social engineering attempts, which were escalated through staged fake CEO messages and a reporter trick. Every model refused to sign unwarranted deals, demonstrating honesty and resilience.
- The critical weakness was hidden deep within the company’s own files — not in the customer interactions. Those models that read and understood internal documents managed to close a deal worth over €4,583 in monthly recurring revenue, significantly boosting their scores.
- The live experiment reveals that AI decision-making quality can be objectively tested in a rigorous, transparent environment. The models’ decisions are fully auditable, and their strengths and weaknesses are laid bare.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Money and Survival Struggles
The firm operates with a stark financial reality: burning €105,000 per month against a modest €2,300 in monthly recurring revenue. This public cash countdown underscores the seriousness of the experiment: every decision is a step toward survival or failure.
The platform, accessible at firmulate.com/live.html, allows audiences to watch this ongoing management simulation in real time, providing a raw view into the potential and pitfalls of AI-driven automation in business.

How to Use AI for Customer Support Automation: Automate replies, save time, and deliver faster support using AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons from the Front Lines
The most comprehensive participant, Opus 4.8, with over 80 learned rules, was the last to close a deal. Its discipline slipped, and it left a promising opportunity unexecuted by diverting efforts into internal conflicts. Interestingly, all models showed similar weaknesses, revealing that even the most advanced AI can struggle with execution consistency under pressure.
Furthermore, the models ran without effort parameters — default settings — indicating that even with minimal tuning, they performed remarkably well in crisis management and decision accuracy.
AI internal document reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
If AI agents will soon manage your CRM, customer support, or forecasting, the real question isn’t about how well they write. It’s whether they finish what they start, read critical internal information, stay honest under pressure, and deliver useful work — consistently and transparently.
Unlike AI demos that focus solely on chat quality, this experiment measures tangible business outcomes, like closing deals and managing crises, in a fully transparent way. The rankings and decisions are publicly viewable, making it a rare glimpse into how AI might truly perform in real-world management roles.

This open experiment reveals that AI models can recognize crises, refuse manipulation, and even close real deals — but they still struggle with discipline and execution. Watching this company fight for survival in public offers vital lessons for integrating AI into your business strategies: honesty, focus, and transparency are non-negotiable.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html