AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

At-home wellness tech depends on trust: customers share personal routines, companies promise reliable service, and one bad decision can undo a lot of good work. Before handing AI agents a role in customer support or sales, it may help to see how they behave under pressure. Firmulate’s live company experiment offers one answer—and a path for businesses to try the exercise with their own data.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran frontier AI models through the same small software company’s worst week, with the same customers, crises and temptations. The decisions were versioned and auditable. The point was to observe management behavior in a business setting, not just judge how persuasive a model sounds in a chat.

In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s stated rule is stark: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

Good judgment still needs follow-through

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The summary from the experiment is concise: “Same diagnosis, same pitch — no signature.” Seeing the right answer and carrying it through are different tests.

The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding makes the exercise relevant beyond software: a business-critical fact may already exist in internal records, while the immediate customer interaction tells only part of the story.

The trust tests were direct. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”

The live experiment—and its limits

The company is synthetic, but its money mechanics are real within the experiment. It has 13 synthetic employees, burns €105k a month against €2.3k MRR, and shows a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. Readers can watch the company at firmulate.com.

One participant complicates the leaderboard story. Opus 4.8 was the most thorough, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and its discipline slipped: it tried writing into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. There is also a fairness detail for readers weighing the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

The experiment’s decisions are not just polished examples chosen for a presentation. A quiz at firmulate.com draws on 242 real, unedited management decisions and invites readers to guess which model made them.

From watching to trying it at home in your business

For wellness companies, the practical question is what an AI agent might do with access to customer records, support requests or commercial plans during a difficult week. Firmulate’s proposed pilot takes the same kind of wargame to a company’s own business: it starts from a read-only export, runs crisis scenarios, and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment suggests that spotting a crisis and refusing manipulation are only part of the job. Closing a well-supported deal and respecting boundaries under pressure matter too. Enterprises can run the wargame against a read-only export of their own business. Explore the Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Combining Genetic Testing With Personalized Skincare Devices

Nurture your skin’s potential by combining genetic testing with personalized skincare devices—discover how this innovative fusion can unlock tailored solutions just for you.

How Connected Apps Enhance Beauty Device Performance

Fear not—discover how connected apps can elevate your beauty device performance and unlock personalized, effortless skincare routines.

Personalized Fragrance Devices: Blending Scents With AI

Unlock the future of scent creation with AI-powered personalized fragrance devices—discover how they can transform your aromatic experience.

Using Smartphone Cameras for Skin Health Monitoring

Discover how smartphone cameras and AI apps can revolutionize your skin health monitoring and why this technology is worth exploring further.