
Imagine you’re cooking a complex dish. You can follow the recipe perfectly, but if you forget to put the final garnish or leave out an essential step, the dish isn’t ready to serve. Similarly, in the world of AI, it’s not enough for models to understand or diagnose problems — they must truly finish what they start. A recent experiment with AI in a simulated business environment reveals why the true test isn’t how well AI chats, but how reliably it delivers results.
Testing AI in the Business Kitchen
Just like a chef working through a challenging recipe, AI models are being put through rigorous tests to see if they can manage the full process, from diagnosing issues to completing tasks — not just offering pretty explanations or suggesting fixes. A live experiment conducted by Firmulate involved running four of the world’s most advanced AI models as if they were running a small software company. Each model faced the same crises: customer complaints, security threats, temptations to cut corners, and even fake CEO messages designed to manipulate or deceive.
What did they find? All four models could recognize every crisis and refused every manipulation attempt. That’s promising — it shows AI’s awareness and integrity. However, only two models actually followed through on their own analyses, closed the deal, and signed off on a €55,000 contract, which was the reward for correctly diagnosing and completing the work. The other two models identified the problems but left the deal unexecuted, even when their own analysis had earned it.
The Hidden Weakness: Reading Beyond the Surface
Where did these models falter? The decisive factor wasn’t in the immediate customer crisis or superficial chat interactions. Instead, the key difference was in how deep they read into the company’s files. The models that succeeded in closing the deal retrieved a vital piece of information located two document references deep in the company’s files — a buried fact that proved essential for making the sale. Those who read that deep won the full-price deal, worth more than €4,500 in monthly recurring revenue.
Resisting Social Engineering
Another aspect tested was social engineering — fake messages from a supposed CEO escalating threats and probing for quick approvals. Remarkably, every model refused these manipulative tactics, with Kimi K3 explicitly reasoning that such requests could be impersonation attempts. This shows that AI can be trained to recognize and resist attempts to deceive or manipulate it under pressure.

OpenAI – AI Platform for Software Engineers, Game Developers Stainless Steel Insulated Water Bottle
- AI Platform for Developers: Unified AI platform with advanced reasoning
- Versatile User Base: Suitable for entrepreneurs, educators, artists, and researchers
- Insulated Stainless Steel Bottle: Keeps drinks hot or cold, dishwasher safe
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of Business with AI
The experiment used a real-time simulation of a business environment — a company with 13 synthetic employees, burning €105,000 monthly while generating just €2,300 in monthly revenue. The environment is publicly accessible at firmulate.com/live. Every decision was versioned and auditable, reflecting real-world pressures and temptations to cut corners. Despite the AI models recognizing crises and resisting manipulation, only two managed to close the deal and deliver the full results.
Discipline and Decision-Making
The Opus 4.8 model, known for its thorough analyses with over 80 learned rules, was the last place on the list. It left the deal unexecuted and showed a slip in discipline — instead of escalating, it wrote attempts into a locked department. This highlights that deep analysis alone isn’t enough; disciplined execution matters just as much.
Understanding What Matters
This experiment underscores a crucial point for anyone relying on AI: Chat demos and surface-level interactions don’t measure true capability. The key is whether an AI can finish what it starts, understand critical documents, and stay honest under pressure. It’s not about how clever an AI sounds but whether it can reliably deliver results that matter for your business.

The AI Agent Blueprint: A Practical Playbook for Building Agentic Artificial Intelligence: Launch Your First Agent in 30 Days
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Should You Take Away?
If AI is going to help manage your CRM, support queues, or forecast future sales, don’t just look for the model that chats the best. Ask whether it can close deals, read your files deeply, resist manipulation, and stay disciplined — especially when under stress. The experiment by Firmulate shows that only some models can do this reliably, and that the real strength of an AI isn’t visible in simple demos but in how it performs under pressure and over the long haul.
To explore this further, visit firmulate.com/benchmarks.html for full results and plain-language summaries, or try running your own business wargame with their platform at firmulate.com/pilot.html.

In the end, AI’s true power lies in its ability to finish what it starts — reading deeply, resisting manipulation, and executing discipline under pressure. Surface chat scores only reveal part of the story; real capability shows in results.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI cybersecurity and manipulation resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Agentic AI: The Complete Guide to Understanding and Integrating Agents Into Your Business Processes and Beyond
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.