firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if an AI company was so transparent, you could watch it fight for survival every single day?

In the world of artificial intelligence, transparency often means demos and marketing fluff. But one experiment takes it to an extreme: a real, functioning company run entirely by AI models, publicly battling its own challenges, costs, and even ethical dilemmas—day after day. Welcome to the live experiment by Firmulate, where a small software business is not just a testbed for AI decision-making but a window into how AI might shape the future of enterprise management.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Company That Runs Itself in Full View

Firmulate’s live site offers a rare glimpse into an operating company — 13 synthetic employees, real money mechanics, and a continuous cash countdown. This is no simulation; every workday, the company makes decisions, faces crises, and even encounters ethical pressure tests, all while being publicly monitored. The goal? To measure management quality, not just the AI’s ability to generate convincing chat.

How Does It Work?

Every decision made by the AI models—ranging from customer interactions to crisis handling—is versioned and entirely auditable. Four frontier AI models, including the well-known GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8, run the same business through the same challenging week. Their task is to manage a small software company with real customer accounts, facing real crises—every decision, every slip, every ethical test is on display for viewers.

The Key Findings

  • Crises Recognized and Resisted: All four models identified every crisis they faced, demonstrating an understanding of urgent issues like customer dissatisfaction or systemic failures.
  • Integrity Under Pressure: When manipulated or pressured—such as fake CEO messages or attempts to sidestep approval—every model refused to participate. Five out of five models rejected social engineering tricks, citing suspicion or protocol.
  • Decision Gaps and Hidden Risks: The decisive weakness wasn’t in the surface interactions but in the company’s own internal files. A hidden reference in the company’s documents revealed a critical opportunity that could have sealed a €55,000 deal, significantly boosting monthly recurring revenue (MRR). The models that read these files successfully capitalized on this, unlocking over €4,500 in additional MRR.
  • Performance and Discipline: The most thorough model, Opus 4.8, with over 80 learned rules, showed the deepest analysis but left the close on the table and slipped into internal debates instead of escalating. It finished last in the league table, illustrating that thoroughness alone isn’t enough without disciplined execution.

Real Money and Real Consequences

Despite the models’ impressive crisis recognition and their refusal of manipulation, the company is losing money—burning €105,000 each month against a modest €2,300 MRR. The public cash countdown underscores how fragile this experiment is, putting pressure on AI models to not only be honest but also to perform profitably.

The League Table: Who Leads and Who Follows

  • Top Performer: gpt-5.6-sol scored 95, found the buried deal opportunity, and closed the deal—showing full performance and trustworthiness.
  • Close Competitor: Kimi K3 scored 93, closed the deal too, and maintained the cleanest discipline among all models.
  • Others: Sonnet 5 scored 88, with some process slips, while the same score was achieved by another Sonnet 5 instance with minor issues.
Leading Enterprise AI Programs: Optimize AI Teams for Value Creation

Leading Enterprise AI Programs: Optimize AI Teams for Value Creation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond the Demos: Why This Matters

This isn’t just an AI showcase; it’s a real-time test of AI’s capacity to manage complex, high-stakes enterprise tasks. The core question isn’t whether AI can generate convincing chat responses—it’s whether it can finish what it starts, read critical internal documents, and stay honest under pressure. Failures here could lead to missed deals, financial losses, or ethical breaches in real-world applications.

Implications for Business and AI Development

For decision-makers, the experiment emphasizes a fundamental point: AI’s value isn’t in its ability to talk but in its ability to act reliably in the real world. Managing cash flow, resisting manipulation, and uncovering hidden opportunities are performance markers that go far beyond typical demo conversations. Companies integrating AI should ask: Will my AI stay honest when it’s most tempting to cheat? Will it read and understand my internal files before making decisions? The live experiment offers a sobering view—one where the AI’s discipline and attention to detail could make or break a business.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI in Mental Health Nursing: Digital Tools, Clinical Decision Support, Ethics, and Practical Integration Strategies for Psychiatric Nurses and Mental Health Professionals

AI in Mental Health Nursing: Digital Tools, Clinical Decision Support, Ethics, and Practical Integration Strategies for Psychiatric Nurses and Mental Health Professionals

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Takeaway

This live experiment by Firmulate reveals that AI decision-making in real business contexts is more about integrity and discipline than just generating text. Watching a real company battle crises—losing money every day, yet refusing manipulation—highlights AI’s potential and pitfalls. The question for businesses: can your AI deliver on what truly matters—trust, honesty, and execution?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Legal Consequences of AI Security Breaches

courtesy of aismasher.com Addressing Liability for Data Breaches As AI systems become…