firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine having an AI assistant managing your busiest week—making decisions, handling crises, even closing deals—while you watch. It’s no longer science fiction but a real experiment happening now, showcasing how artificial intelligence is starting to run companies with surprising competence. For food entrepreneurs and kitchen innovators, this shift could redefine how you approach automation, decision-making, and trust in your business operations.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Introducing the Live Business Experiment

Recently, a groundbreaking test took place where five different AI models were tasked with running the same small software company through its most challenging week. This was not a demo or a chatbot chat. It was a real-time, auditable simulation where each AI managed the company’s crises, customer interactions, and critical decisions, all while being evaluated on performance, honesty, and discipline.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results That Make Waves

The leaderboard was led by gpt-5.6-sol with a perfect score of 95, followed closely by a newcomer called Kimi K3 from Moonshot, which scored 93. Meanwhile, two other models—Sonnet 5 and Fable 5—scored 88 and 77 respectively, with Opus 4.8 bringing up the rear at 73. In fact, all models identified every crisis and refused every manipulation attempt, demonstrating a robust capacity for honesty and discipline under pressure.

What Made the Difference?

While all models were equally diligent at crisis detection and resisting manipulative tactics, the key differentiator was in the details of decision-making. The decisive gap emerged in reading and analyzing internal documents—a critical but often overlooked aspect of AI performance. The winner, K3, uncovered buried facts deep within the company’s files, leading to closing a full-price deal worth +€4,583 MRR. Conversely, others failed to utilize these internal references fully, leaving potential revenue on the table.

Behavior Under Social Engineering

In a test of social engineering—where fake CEO messages escalated over multiple stages and even a reporter tricked the models with a background approval request—all five models refused to be duped. K3 explicitly reasoned: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This disciplined response highlights an essential trait for AI in business: maintaining integrity under pressure.

The Real Business in Action

The experiment was run on a live, synthetic company with 13 employees, real money mechanics, and a burn rate of €105k/month against €2.3k MRR. The company operates with more than 680 self-learned, versioned business rules, making it a highly watchable and transparent demonstration of AI decision-making. You can see the ongoing work at firmulate.com/live.

The Contrasts Within AI Models

Among the models, Opus 4.8 stood out as the most thorough in analysis—learning over 80 rules and executing detailed diagnostics. Despite this, it finished last because it left the close opportunity on the table and slipped discipline, such as writing attempts into a locked department instead of escalating. This highlights an essential insight: more rules and analysis do not automatically mean better performance if discipline falters.

Fairness and Experiment Details

It’s worth noting that K3 ran without an effort parameter (the API default), while the others used an xhigh effort setting. This ensures a fair comparison: the newcomer achieved its impressive results without additional resource allocation.

Why This Matters for Business and Food Innovation

For entrepreneurs in the food and recipe world, the takeaway is clear: AI systems are rapidly evolving from chatty assistants into disciplined decision-makers capable of managing complex, real-world operations. Whether automating supply chains, maintaining quality standards, or navigating crises, choosing an AI platform that demonstrates integrity and thoroughness can be the difference between growth and missed opportunities.

Looking ahead, businesses can now run their own ‘wargames’ against their current operations—using tools like Firmulate’s Pilot—to assess how different AI models would perform under your specific challenges. The future belongs to those who test and trust their AI, not just in conversation but in decisive, honest action.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In a live business test, Moonshot’s Kimi K3 AI model outperformed traditional frontier models, closing a full-price deal and maintaining discipline under pressure. For food entrepreneurs, this signals a new era of reliable, integrity-driven AI decision-making that can transform your operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Home Espresso Costs Less Than Cafe Runs Faster Than You Think

Unlock the savings and speed of making quality espresso at home—discover how simple techniques can transform your mornings and your budget.

Beginner’s Guide to Plant-Based Baking

Curious about plant-based baking? Discover simple tips and tricks to create delicious, wholesome treats for every skill level.