AIThis post was created with the assistance of artificial intelligence (AI).

A bakery’s worst day can start with a late delivery, a key customer threatening to leave, or a suspicious message claiming to come from the boss. In food businesses, small decisions can affect trust as quickly as they affect the day’s takings. Firmulate’s live experiment asks what happens when AI is put in charge during a company’s worst week—and whether it can follow through when the right answer calls for action.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same week, different models

Firmulate ran several frontier AI models through the same small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The benchmark’s stated principle is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Knowing what to do is not the same as doing it

All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The gap matters beyond software. A food business might have an AI assistant identify a promising wholesale order or a supplier risk; the practical question is whether it can handle the next step while respecting the business’s rules.

The deal hinged on a detail buried two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to an everyday management challenge: relevant information can sit in records a person or system must actively consult.

Trust under pressure

The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 ran without an effort parameter, using the API default, while the others ran at xhigh—context that matters when comparing the results.

A company you can watch

The experiment is attached to a live, synthetic company with 13 employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

For companies considering AI in customer service, sales or operations, the experiment offers a practical next step: test against your own situation. Firmulate says enterprises can run the wargame using a read-only export of their business, then receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

From watching to a pilot

A model can identify a crisis and resist a trick, yet still miss the action its own analysis supports. Firmulate’s live experiment makes that difference visible; a pilot can put a company’s own data and playbooks to the test. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mountain Dew Will Release a 5-Cent Limited-Edition Soda — How to Get It

Mountain Dew will release a limited-edition soda at 5 cents, available through a special promotion. Here’s how to get it and what it means.

James Beard Foundation Awards 2026: Winners, News, and Updates

The James Beard Foundation revealed the winners of the 2026 awards at its June 15 ceremony in Chicago, honoring top chefs, restaurants, and media in American food.

KitchenAid Artisan vs KitchenAid Classic: Full Comparison

Compare the KitchenAid Artisan and Classic Series 4.5 Qt mixers to find the best fit for your baking needs. Durable, versatile, and reliable options explained.

Build vs Buy a Prebuilt AI Workstation

Struggling to choose? Discover whether building or buying a prebuilt AI workstation makes more sense today with real costs, performance, and support insights.