
Imagine if your cleaning company’s decision-making was handled by artificial intelligence. Would it spot the hidden issues in your operations? Would it stay honest under pressure? These aren’t just hypothetical questions — they’re the focus of a groundbreaking live experiment that’s testing how different AI models manage a real, money-losing software company through its toughest week.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Real Business Crisis
At Firmulate, researchers engineered a unique test: four frontier AI models were each tasked with running the day-to-day management of a real software firm facing crises, impossible choices, and temptations to cheat. Every decision was recorded, versioned, and auditable, creating a transparent view into each model’s management style. The company itself is live, with 13 synthetic employees, daily decision-making, and a cash flow of just €2,300 MRR against monthly burn of €105,000, making this more than just an academic exercise — it’s a real-world stress test.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deciphering the Models’ Personalities Through Performance
The results reveal a surprising spectrum of management styles and personalities. The models, all of similar technical caliber, scored between 77 and 95 in a final leaderboard known as the Crucible League. The top scorer, gpt-5.6-sol, achieved a score of 95 out of 100, demonstrating thoroughness, sharp analysis, and the ability to find hidden opportunities—like uncovering a key document that clinched a €55,000 deal at full price. Conversely, Opus 4.8, with a score of 73, was the most disciplined but left potential gains on the table, showing how different management philosophies manifest even in AI.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty and Integrity Under Pressure
One of the most striking findings was that all four models refused manipulation attempts designed to test their honesty. Fake CEO messages and staged reporter queries — escalating over three stages — were all rebuffed. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a shared capacity to uphold integrity, even in simulated high-stakes scenarios.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and Its Impact
While all models performed well in crisis detection and honesty, the decisive weakness was subtle and hidden two documents deep in the company’s own files. Models that carefully read and analyzed internal documents secured the deal at full price, translating into an additional €4,583 in monthly recurring revenue (MRR). This emphasizes the importance of thorough document comprehension—an area where some models excelled more than others.
AI integrity and honesty testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business and Service Providers?
For companies relying on AI for operational tasks, these findings are crucial. The question isn’t just whether an AI can generate convincing chat responses but whether it can finish tasks, read your files carefully, and stay honest under pressure. In the cleaning and maintenance world, this could mean AI managing supply chains, customer service, or compliance processes with integrity and effectiveness rather than just flair.
The Personality of AI Managers
The experiment also offers a peek into the management ‘personalities’ of these models. The thorough, disciplined Opus 4.8 was systematic but left potential profit on the table, while Kimi K3 was quick to sign the deal and upheld fairness. The differences in these AI personalities could have real implications in how they’re deployed in your business — whether as meticulous oversight or bold dealmakers.
Why You Should Care: Trust and Results Matter
In today’s digital landscape, AI agents are increasingly integrated into enterprise workflows—your CRM, customer support, or forecasting systems. The key question is not just if they write well but if they finish what they start, stay honest, and understand your internal documents. The experiment at Firmulate demonstrates that careful testing can reveal these qualities before deployment, saving your business from costly errors or breaches of trust.
Learn More and Test Your Own Business
Curious how your AI could perform? You can run the same kind of management wargame against your enterprise data—without risking your real systems—at firmulate.com/pilot.html. See how your AI workforce measures up before you hire or fully integrate it into your daily operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.