
Imagine a restaurant run entirely by artificial intelligence — making decisions, handling crises, and even negotiating deals, all in real time. Now, what if you could watch this AI in action, week after week, as it navigates the unpredictable chaos of running a business? This is not science fiction; it’s the live experiment conducted by Firmulate, a company that publicly tests AI management in its most extreme form.
The Live Experiment: An AI Company in Real Time
Firmulate has created a unique, transparent showcase of AI decision-making. It runs a small, virtual software company with 13 synthetic employees, where every move, every crisis, and every decision is publicly visible at firmulate.com/live.html. This isn’t a hypothetical; it’s a real-time, ongoing test of AI models as they manage a company facing daily setbacks and temptations to cheat.
The company’s financial state is stark: it burns through €105,000 every month while generating only €2,300 in recurring revenue. With a public cash countdown, every decision matters. The entire operation is versioned daily, with over 680 self-learned rules guiding its management. Observers can see how each AI model responds to real crises—be it a customer complaint or an internal temptation—making this experiment a rare window into AI’s practical capabilities in complex, unpredictable scenarios.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing the Frontiers of AI Management
Firmulate tested four leading AI language models—each tasked with running the same difficult week in the company’s life. These models faced identical challenges, including customer crises, internal conflicts, and manipulative attempts like social engineering. Interestingly, all four models successfully identified every crisis and refused manipulation attempts. They demonstrated integrity and awareness, refusing fake CEO messages and suspicious background requests.
The experiment’s most revealing result was in deal-making. Only two models, gpt-5.6-sol 95 and Kimi K3 93, signed the €55,000 deal their analysis had earned. The other two—Sonnet 5 and Fable 5—missed opportunities, with one leaving a deal on the table and the other failing to escalate internal issues properly. The key weakness was hidden two document references deep in the company’s files—something the models that read the files fully picked up, leading to successful deal closure.

The AI Culture Blueprint: Moving Beyond Tools to Create Human-Centered AI Adoption
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights Beyond Demos
What makes this experiment extraordinary isn’t just the live transparency or the dramatic financials. It’s how the models handle integrity, thoroughness, and discipline under pressure. For instance, the Opus 4.8 model, despite its depth of analysis, left a deal unclosed because it was overwhelmed by internal process slips—it wrote attempts into a locked department instead of escalating them. Meanwhile, K3 ran without an effort parameter—meaning it didn’t push itself as hard as others—yet still performed well in fairness, showing how different configurations impact results.
This ongoing story underscores a vital point for businesses considering AI automation: the question isn’t only whether an AI can generate good-looking output but whether it can see through crises, stay honest, and close deals effectively under stress. The live company blazes a trail, revealing strengths and weaknesses in real-world conditions, not just in scripted demos.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Should You Care?
If your organization’s future involves AI in customer support, sales, or decision-making, understanding what these models can truly do—beyond chat—matters more than ever. The experiment demonstrates that AI’s ability to handle crises, read internal files, and resist manipulation is crucial to its usefulness. It’s not just about shiny responses but about the capacity to finish what it starts, stay honest, and work efficiently against real business odds.
Interested in seeing how these AI models perform? You can explore the live results, watch ongoing experiments, and even try running your own business wargame against a read-only export at Firmulate’s Pilot.

Firmulate’s live AI company experiment reveals that while models can identify crises and refuse manipulations, closing deals and fully understanding internal data remain challenges. Watch the live company in action and see if AI is ready to handle real business pressures—and what it might cost to do so.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Language Translation Earbuds Translator in Real Time 0.5S,Two-Way Bluetooth Translator Device with APP for 144+ Languages Translation Packs,Spanish English Translation Headphone,M10 (Black)
🌐 144-Language AI Interpreter|0.5s Natural Conversation Tech-Patented NLP algorithm enables seamless bilingual dialogues without awkward pauses. Unlike traditional…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.