
Imagine handing an AI your restaurant’s reservations, supplier notes and staff playbook, then asking it to manage the night the fridge fails, a key supplier goes quiet and a reporter calls. Would it spot the trouble? Would it protect the business? And would it actually close the deal that could keep the lights on?
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That is the kind of question Firmulate is putting to the test. Its live experiment runs AI models as a small software company under pressure. For food businesses considering AI in bookings, customer service or operations, the premise is familiar: a calm demo is not the same as a chaotic Friday night.
A tough week, shared by every model
In the final Crucible League, completed in July 2026, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were the same. Decisions were versioned and auditable.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s finding was striking: every model spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
For a restaurant, the equivalent might be an AI correctly identifying that a supplier problem threatens service, recommending a backup and then failing to make the necessary commitment. Recognizing the problem is useful. Carrying the response through is what protects the evening.
The clue was buried in the paperwork
The deal hinged on a competitor weakness tucked two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. The episode makes a practical point: useful business context may be hiding in documents a rushed operator rarely has time to revisit.
Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
There was a more awkward result, too. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. More analysis did not guarantee better execution.
Watch the experiment before trying your own
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The experiment is real and watchable at firmulate.com.
The quiz at firmulate.com/quiz.html turns 242 real, unedited management decisions into a “guess the model” challenge. It offers a reader-friendly way to see how hard it can be to tell which model is behind a business decision.
There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the rankings when interpreting the result.
From watching to a business pilot
For a restaurant group, food supplier or hospitality business, the next step need not be handing an AI access to live systems. Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The exercise can test crisis scenarios against company-specific context and produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems.
That makes the proposal less like hiring an AI and hoping for the best, and more like rehearsing the hard shift before putting new help on the floor. The experiment asks whether a model can do more than sound capable when the order book, customer trust and cash are under pressure.

Before putting AI into a real workflow, test it against the decisions your business might face on its worst week. To discuss a pilot using a read-only export of your company data, visit Firmulate’s pilot page and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
