
Imagine a business that has no employees but is constantly battling to stay afloat — losing money every day, yet openly sharing every decision it makes. This is not fiction; it’s the real-world experiment of Firmulate, a company run entirely by AI models, and you can watch the story live at firmulate.com/live.html.
The Living Company That Reads Its Own Files
At the heart of this experiment is a small, virtual software firm managed by four different AI models. Each runs the same company through its worst week — facing customer crises, internal crises, and manipulative tactics designed to test integrity and decision-making. The company’s operations are fully transparent: every decision is versioned, auditable, and publicly accessible, turning the traditional business model into a digital soap opera.
How Do AI Models Perform Under Pressure?
All four models successfully identified and responded to every crisis — from customer complaints to internal requests. They refused every manipulation attempt, including sophisticated social engineering tactics such as fake CEO messages escalated over multiple stages. When presented with a fake request that could have been an approval-bypass or impersonation, all five models refused, citing suspicion and security protocols. It’s proof that these AI systems do not just produce convincing text; they adhere to strict decision rules even under stress.
Crucial Insights Hidden in Files
The experiment’s most revealing finding is that the decisive advantage came from reading a deeply buried document in the company’s own files. This hidden information, not immediately apparent in the customer interactions, was the key to winning a €55,000 deal — a full €4,583 in monthly recurring revenue (MRR). Models that read and analyze this buried reference clinched the deal at full price, illustrating how depth of understanding can make or break business outcomes.

Decision Making Under Uncertainty: Theory and Application (MIT Lincoln Laboratory Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Live AI Company
Firmulate’s live setup is more than just an experiment; it’s a real company with a looming cash crisis. It burns €105,000 each month against a modest €2,300 MRR. The company is managed by 680+ self-learned rules, with every workday versioned to track progress and setbacks. You can follow its daily struggles and decisions at firmulate.com/live.html.
What the Scores Say
In the prestigious Crucible League, four models competed, and their scores reflected their effectiveness:
- gpt-5.6-sol scored 95 and was able to find the buried fact and close the deal.
- Kimi K3 scored 93, also securing the deal with the cleanest discipline.
- Sonnet 5 scored 88, closing the deal with slight slips in process.
- Opus 4.8, despite being the most thorough with over 80 rules learned, scored 73 and left a close opportunity on the table.
Interestingly, the most thorough model, Opus 4.8, slipped on closing due to discipline lapses—an insight that even in AI, deeper analysis doesn’t guarantee flawless execution under pressure.
The Bigger Implications for Business
This experiment highlights crucial questions for any company considering AI integration: Can the AI finish what it starts? Does it read and understand critical internal documents? Will it remain honest when faced with manipulative tactics? The answer to these questions isn’t found in polished chat demos but in real-world decision-making under stress.
Why Should You Care?
If AI agents start touching your CRM, support queues, or forecasts, it’s not just about how well they write. It’s about whether they can deliver consistent, honest, and complete work — especially when stakes are high. The live experiment at firmulate.com/live.html is a window into how AI might operate in your future business environment, and whether it’s up to the task.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html