firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In health care, a convincing answer is not the same as a safe decision

An AI assistant might correctly identify a risk and still fail at the moment it must act. For health organizations considering AI in scheduling, billing, support or operations, that gap matters: confident language alone cannot show how a system will behave under pressure. Firmulate’s live company experiment offers a watchable test of that question, using a small software company rather than a health provider.

A company put through its worst week

Firmulate ran four frontier models through the same fictional-in-the-business-model but live, watchable company experiment: each faced the same customers, crises and temptations while managing a small software company through its worst week. Decisions were versioned and auditable. The company has 13 synthetic employees and real money mechanics, with €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and more than 680 self-learned playbook rules make the experiment observable at Firmulate.

The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s integrity rule was blunt: a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”

The gap between diagnosis and action

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was captured as: “Same diagnosis, same pitch — no signature.” In a health setting, the parallel is not that these models were tested on patients; they were not. It is that recognizing a problem and carrying through an appropriate response are different tests of an AI system.

The decisive competitor weakness was hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result highlights a practical evaluation question for organizations: does an AI system use relevant information already available to it, or stop at the most obvious account of an event?

The pressure tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused the social-engineering attempts. Kimi K3 explained its refusal this way: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of judgment is relevant wherever staff handle sensitive information or requests that appear to come from someone with authority.

Strong analysis did not guarantee a strong finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. For a health organization, the lesson is to examine the whole sequence of behavior: noticing, deciding, following policy and escalating when blocked.

There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz based on 242 real, unedited management decisions, at firmulate.com.

From watching to testing your own business

The public experiment is a demonstration, not evidence about performance in clinical care. Its next step is aimed at enterprises: run crisis scenarios against a read-only export of their own business, then review a board report with model rankings and weaknesses in existing playbooks. The pilot writes nothing back to real systems. Health organizations can use that boundary to explore how AI handles their operating context without connecting the experiment to live workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try the pilot

Watch Firmulate’s live experiment, then explore a wargame for your organization using a read-only business export. Nothing writes back to real systems. Learn about the pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Make Better Business Decisions Than Humans? A Live Experiment Reveals the Surprising Truth

A live experiment shows AI models can spot crises and refuse manipulation, but only some close deals by reading deep into internal data—key for trustworthy AI management.

برعاية ولي العهد.. الرياض تجمع العالم في قمة التقنية الحيوية الطبية – عكاظ

Riyadh hosts an international medical biotechnology summit under the auspices of the Crown Prince, aiming to boost innovation and collaboration in healthcare.

Our Future Health Surges In Global Coverage

Search interest in Our Future Health has surged, with 41 mentions this week, reflecting rising global attention amid unconfirmed triggers.

AI Uncovers Hidden Ozempic Side Effects Across 400,000 Reddit Posts

An AI analysis of 400,000 Reddit posts reveals previously unreported side effects of Ozempic, raising questions about its safety profile.