AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your pool business hold up through a week of cancellations, a price squeeze and a competitor’s offer?

For a company in the pools, patio and water lifestyle trade, the hard days rarely arrive one at a time. Customers change plans, cash keeps moving and a rival may spot an opportunity buried in your own paperwork. Firmulate’s live experiment puts an AI-run company through that kind of pressure, with real money mechanics and decisions people can watch. The next step is to try the exercise with your own business.

A live company, not a chat demonstration

Firmulate describes itself as an AI company emulator. Its live company has 13 synthetic employees, a public cash countdown and more than 680 self-learned playbook rules. It spends €105,000 a month while bringing in €2,300 in monthly recurring revenue. Workdays are versioned, so the choices can be followed over time at firmulate.com.

The final Crucible League, in July 2026, put five models through the same small software company’s worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules treated partial progress as progress, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Spotting trouble wasn’t the hard part

Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The finding is stark: “Same diagnosis, same pitch — no signature.” An AI can identify a sound move and still leave it undone.

The winning detail was not in the customer event. The decisive competitor weakness was two document references deep in the company’s own files. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. For a pool builder or patio retailer, the parallel is practical: the clue that changes a negotiation may already be in a customer record, quote history or internal playbook.

Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong analysis still needs follow-through

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet it finished last. It left the deal on the table and slipped on discipline, attempting writes in a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. More analysis alone did not guarantee a better business outcome.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league is a snapshot of this experiment, not a promise about how any model will perform in every company.

Visitors can also try a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com. The decisions offer a closer look at the difference between recognizing a problem and managing it.

From watching to trying it on your business

The proposed enterprise pilot takes the exercise from the live company to a participant’s own business. It starts from a read-only export, then runs crisis scenarios against a digital twin and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise relevant to businesses weighing AI agents for customer service, sales or planning: they can examine behavior under pressure before connecting an AI to operational tools.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decisions before they reach your systems

Firmulate’s experiment suggests that crisis recognition and sound analysis do not always lead to a completed deal. A pilot can put your company’s scenarios and playbooks under similar pressure using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Langelier Saturation Index: Balancing Scale and Corrosion

Maintaining optimal water chemistry with the Langelier Saturation Index is crucial for preventing scale and corrosion; discover how to balance your system effectively.

How Saltwater Pools Still Depend on Smart Chlorine Management

Of course, saltwater pools still rely on smart chlorine management to stay safe and clear, but the key details might surprise you.

Dilution & Partial Drains: Resetting Water Chemistry Safely

Just by carefully using dilution and partial drains, you can reset your water chemistry safely—learn how to do it right and keep your system healthy.

The Chlorine Demand Clues Most People Miss Until It’s Too Late

Learning to spot early chlorine demand signs can save your pool—discover the crucial clues most owners overlook before it’s too late.