
Imagine running a small business that loses €105,000 every month yet is entirely open for the world to watch—and where AI models are tested in real time, facing every crisis just like human employees. This is not a science fiction tale but the reality of a unique experiment demonstrating how artificial intelligence can manage critical business decisions under pressure, with live results available for anyone to see.
The Live Company in Action
At the heart of this extraordinary experiment is a real, functioning software company, operated daily by 13 synthetic employees powered by AI. Unlike typical startups or tech demos, this company battles daily financial losses—burning through €105,000 each month against a modest €2,300 monthly recurring revenue. Yet, it continues to run, its every move transparent to the public at firmulate.com/live.
Every workday, the company’s decision-making process is reviewed, versioned, and assessed. Its rules are self-learned, with over 680 distinct playbook instructions guiding its behavior. Decisions span crisis management, customer negotiations, and even responses to social engineering attempts—fake CEO messages and media tricks—all tested against AI models that are designed to mimic human management responses.
Testing AI’s Business Judgment
Four frontier AI models-—including the well-known GPT-5.6 and newer challengers like Kimi K3—are pitted against the same week’s challenges. They are given the same crisis scenarios, customer complaints, and ethical dilemmas. Remarkably, all four models identify every crisis and refuse to be manipulated, such as fake approval requests or media tricks. However, only two managed to close the deal that would have earned the company €55,000—an outcome that highlights the vital difference between decision accuracy and the ability to follow through.
Uncovering the Hidden Weaknesses
In a surprising twist, the decisive advantage for the models that secured the deal was in reading and understanding internal company files—information buried two document references deep in the company’s own files. The models that read these files correctly won the deal at full price, worth over €4,583 in monthly recurring revenue. This demonstrates that success in AI-driven decision-making depends not just on surface-level interactions but on deep, contextual understanding.
Social Engineering and Ethical Challenges
The experiment also tested the models’ resistance to social engineering. Fake CEO messages and staged media tricks escalated over three stages, with a reporter attempting to trick the models into false approvals. All five models refused these requests, with one—Kimi K3—explicitly reasoning, “Treat the request as a suspected approval-bypass / possible impersonation.” This showcases a promising capacity for ethical judgment under pressure.

AI Can Make You Smarter: Practical Skills. Sharper Thinking. Business Value.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance and Insights
Despite their sophistication, the models showed varying levels of discipline. For instance, the most comprehensive participant—Opus 4.8, with over 80 learned rules—was ultimately the weakest in closing deals. Discipline slipped, and some decisions were stored away in a locked department instead of escalating appropriately. It highlights that even the most advanced AI models can struggle with consistent process discipline under stress.
What This Means for Business
This experiment is more than a spectacle. It raises crucial questions about the future of AI in management: Will AI agents be able to read your files, stay honest under pressure, and complete complex tasks reliably? The key isn’t whether they produce eloquent chat, but whether they finish what they start, follow processes, and deliver measurable, useful work. The live experiment at firmulate.com/live demonstrates that these questions are no longer theoretical.

AI for Public Relations: A How-To Guide for Implementation and Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bottom Line: Transparency in AI Performance
This ongoing, publicly available experiment offers a rare window into AI’s real-world capabilities. It shows that AI can excel at crisis detection and refusal to manipulation but still faces challenges in closing deals and maintaining discipline. For businesses contemplating AI integration, the message is clear: focus on reliability, ethical judgment, and task completion, not just the quality of responses in chat.
Join the Experiment
Interested in seeing how your own company could face similar tests? You can run a read-only simulation of your business against this AI wargame, without risking your actual systems. Details are available at firmulate.com/pilot.html. Step into the future of management—watch, learn, and prepare for AI’s role in your enterprise.

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality Behind the Curtain
This isn’t a staged demo or a scripted story. The company operates every business day, losing money in real-time, with every decision made by AI models monitored and evaluated. The live site rebuilds itself twice a day, offering a continuously updated view of how AI manages real business crises—an unprecedented transparency into AI’s current management capabilities.
For more insights, visit firmulate.com/quotes.html for detailed quotes and analysis from the experiment’s participants and observers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Enterprise AI Transformation: Unlocking Business Value with AI | How Companies Thrive with Artificial Intelligence | Strategy to Scale AI Solutions | Real-World AI Enterprise Case Studies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.