
Imagine if your cleaning company’s management decisions were made by AI—decisions that could either save thousands or cost you millions. What if you could test these AI managers before trusting them with your daily operations? Welcome to the live experiment from Firmulate, where real AI models run a small software company through its worst week, revealing not just what they decide but how they think. For businesses that rely on trust, precision, and integrity—like cleaning and maintenance services—understanding AI decision-making isn’t just tech talk; it’s a business survival skill.
The Experiment: Putting AI to the Test in a High-Stakes Environment
Firmulate’s latest showcase pushes four frontier AI models—ranging from GPT-5.6 to Sonnet 5—into a simulated crisis week for a real software company. Every decision is real, every crisis authentic, and each AI faces the same challenges: customer issues, internal conflicts, ethical dilemmas, and manipulative tactics. The goal? See if AI can handle complex management tasks with integrity and accuracy.
Each model’s performance was scored based on its ability to detect hidden risks, maintain discipline, and ultimately close a crucial €55,000 deal—an essential revenue boost that could make or break the company’s week. Interestingly, all models identified every crisis and refused manipulative attempts—showing they’re aware of the ethical stakes. Yet, only two models managed to close the deal, earning the full prize without slipping into shortcuts or misconduct.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: The Power of Reading Deeply and Staying Honest
- All four models spotted every crisis and refused every manipulation attempt—an encouraging sign for trustworthiness.
- The decisive factor was what the models read in the company’s internal files. Those that examined these references fully were able to find a critical piece of information buried two documents deep, leading to the successful deal.
- Models that skipped or skimmed the deep references missed the opportunity for full-value sales, costing the company an additional €4,583 in Monthly Recurring Revenue (MRR).
As an affiliate, we earn on qualifying purchases.
Behavior Under Pressure: Ethical Stances and Manipulation Defense
In a social engineering test, the models faced staged fake CEO messages escalating in intensity, plus a reporter’s subtle background question. All five tested models refused to engage in deceptive or manipulative behavior, with Kimi K3 explicitly treating the request as a possible impersonation attempt. This suggests that current AI models are not just technically capable but also ethically cautious when faced with social engineering tricks.
As an affiliate, we earn on qualifying purchases.
The Real Company Behind the Experiment
The company used in the live test is a real, operational software business with 13 synthetic employees, handling real money mechanics—burning €105k monthly against a 2.3k MRR. It’s a public, transparent experiment, with a continuously evolving setup. Every decision, every rule, and every outcome is versioned and observable at firmulate.com/live. This isn’t just a demo; it’s a window into how AI might manage real businesses under stress, day after day.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Cleaning & Maintenance Businesses
For companies in cleaning, floor care, or maintenance services, AI is becoming a tool—not just for scheduling or customer support but for strategic decision-making. The experiment demonstrates that AI can recognize crises, refuse unethical shortcuts, and even uncover hidden risks if it reads deeply into your operational data. Conversely, superficial AI engagement might miss vital details, leading to costly mistakes or lost opportunities.
It’s crucial to assess your AI tools not just on how well they generate chat or handle customer queries but on whether they finish what they start, read your files thoroughly, and stay honest under pressure. The Firmulate experiment shows that trustworthiness and thoroughness are measurable and essential qualities for AI that manages your business.
Join the Wargame: Test Your AI Workforce
Business leaders interested in evaluating their own AI systems can run a similar simulation—called a wargame—against a read-only export of their operations. This allows you to see how your AI would behave under stress without risking your actual systems or data. Learn more and try it yourself at firmulate.com/pilot.html. It’s an invaluable step before trusting AI with your critical business decisions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html