
Imagine a real company operating every workday, making decisions, facing crises, and even losing money — all run by artificial intelligence models. For health-conscious readers, it’s a bit like trusting your health to a new AI-powered app that claims to optimize your wellness but is still in testing. Would you rely on it? Or would you want to see it prove itself first? That’s exactly what a pioneering experiment is doing — revealing how AI models perform in real-world, high-stakes business scenarios, live and unfiltered.
A Business That’s Live, Public, and Under Pressure
At the cutting edge of AI experimentation, a real, functioning software company is operating every workday with a twist: it has no human employees. Instead, 13 synthetic ’employees’ — powered by advanced AI models — manage the company’s decision-making, money mechanics, and crisis responses. This experiment isn’t just theoretical; it’s happening live, and anyone can watch it unfold at firmulate.com/live.
The company is running every weekday, burning through €105,000 in cash against a modest €2,300 monthly recurring revenue. It’s a high-stakes test of AI’s ability to manage real business risks and opportunities, all under the scrutiny of a public audience. Every decision the AI models make is tracked, versioned, and open for review, creating a transparent window into their decision processes.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge: Surviving the Worst Week
The experiment involved running four different frontier AI models through the same simulated crisis week. The models faced the same customers, the same crises, and the same temptations to cut corners or manipulate data. This rigorous setup aimed to assess how reliably these models could perform under pressure.
The results? All four models identified every crisis and refused every attempt at manipulation. Yet, only two of the models managed to close a deal worth €55,000 — the revenue target — based solely on their own diagnosis and pitch. The other two, despite similar analysis, failed to secure the deal and left the opportunity on the table.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: What the Models Read Matters
Digging deeper into why some models succeeded and others didn’t revealed an intriguing insight: the decisive weakness was buried in a company document, not in the immediate customer interactions. Models that read a specific document reference in the company’s own files uncovered that hidden piece of information, leading to the successful deal. The models that missed it lost out on full revenue, costing roughly +€4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering and Deception
In one of the more realistic tests, the models faced fake CEO messages escalating over three stages, plus a reporter’s trick — a simple request for a background yes/no. All five models tested refused to be manipulated, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates AI’s capacity for resisting social engineering tactics, crucial in safeguarding business operations from deception.
As an affiliate, we earn on qualifying purchases.
The Reality of a Money-Losing Company
The live entity is a real software operation with a tough financial reality. It’s losing €105,000 each month against only €2,300 in monthly revenue, with a public cash countdown looming. The setup features over 680 self-learned rules and daily versioned decisions, making it a living, breathing experiment accessible for anyone curious enough to watch at firmulate.com/live.
This experiment isn’t just about AI prowess; it’s about how AI agents behave under pressure, whether they stay honest, and if they can deliver useful outcomes that matter in real business. For those in health and wellness, it emphasizes a vital point: trusting new technology means watching it perform in real situations, not just in polished demos.
What We Learn About AI in Business
Despite the models’ success in crisis detection and resisting deception, not all performed equally. The most thorough participant, Opus 4.8, analyzed deeply with over 80 learned rules but still left deals on the table and slipped in discipline, such as writing attempts into a locked department instead of escalating. The contrast underscores that technical depth alone doesn’t guarantee success — discipline, process adherence, and understanding the hidden details matter.
Why This Matters for Everyone
In health and wellness, adoption of AI tools will soon be inevitable — for diagnostics, patient management, or data analysis. This experiment shows that the critical question isn’t whether AI can generate good-looking outputs but whether it can complete the work, stay honest, and adapt under real-world pressures. It’s a reminder that transparency and testing in controlled yet realistic settings are essential to trusting these systems in sensitive areas like healthcare.
See It in Action and Decide for Yourself
You can witness this experiment live, explore the decisions made, and even try to guess which AI model made which choice through a public quiz at firmulate.com/quiz.html. Or run your own test against your business data with the pilot program, which ensures no real systems are affected.

This live AI business experiment reveals that the true test of AI readiness is not just in generating answers but in completing real tasks honestly under pressure. Watching a real company burn money daily, managed solely by AI, offers critical lessons for health and wellness sectors about trust, transparency, and the importance of rigorous testing before full adoption.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html