A bakery’s worst day can start with a late delivery, a key customer threatening to leave, or a suspicious message claiming to come from the boss. In food businesses, small decisions can affect trust as quickly as they affect the day’s takings. Firmulate’s live experiment asks what happens when AI is put in charge during a company’s worst week—and whether it can follow through when the right answer calls for action.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same week, different models
Firmulate ran several frontier AI models through the same small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The benchmark’s stated principle is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Knowing what to do is not the same as doing it
All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The gap matters beyond software. A food business might have an AI assistant identify a promising wholesale order or a supplier risk; the practical question is whether it can handle the next step while respecting the business’s rules.
The deal hinged on a detail buried two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to an everyday management challenge: relevant information can sit in records a person or system must actively consult.
Trust under pressure
The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 ran without an effort parameter, using the API default, while the others ran at xhigh—context that matters when comparing the results.
A company you can watch
The experiment is attached to a live, synthetic company with 13 employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
For companies considering AI in customer service, sales or operations, the experiment offers a practical next step: test against your own situation. Firmulate says enterprises can run the wargame using a read-only export of their business, then receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.
From watching to a pilot
A model can identify a crisis and resist a trick, yet still miss the action its own analysis supports. Firmulate’s live experiment makes that difference visible; a pilot can put a company’s own data and playbooks to the test. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
