AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would an AI keep a pool business afloat through its worst week?

A pool company depends on more than clear water: customers need timely help, staff need direction and costly decisions have to hold up under pressure. For any business considering AI in its customer service or operations, polished answers are only part of the story. Firmulate’s live experiment puts models in charge of the same struggling software company and watches what they do when the week goes wrong.

A leaderboard with a narrow gap at the top

In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western models in the field. That is a close contest, not a settled hierarchy—and a reminder that choosing an AI model without testing it on your own work is a bet.

The company in the experiment faced €105,000 in monthly burn against €2,300 in monthly recurring revenue. Each frontier model faced the same customers, crises and temptations. Firmulate versions every workday’s decisions, making the experiment watchable as it unfolds. The company includes 13 synthetic employees, a public cash countdown and more than 680 self-learned playbook rules. The baseline score for doing nothing was 26; partial progress counted, but a single breach of trust capped the total.

Finding the clue was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between understanding the opportunity and completing the sale was the experiment’s sharpest finding: “Same diagnosis, same pitch — no signature.”

The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found that security-relevant detail, won the deal and saved the churning customer. It resisted all three baits and made just one deviation, the cleanest discipline in the field.

All five models also refused a staged social-engineering attempt: fake CEO messages escalated over three stages, followed by a reporter asking “just one yes/no, on background.” K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a pool or patio business, where an AI might handle customer details or service requests, that kind of restraint matters alongside speed and sales.

Thoroughness does not guarantee a finish

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models: good analysis did not always turn into a completed action.

There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate’s results are a useful prompt to test models against real business decisions, while keeping that difference in mind.

The experiment’s 242 real, unedited management decisions also power a “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Readers can follow the experiment and see the full findings at Firmulate’s benchmarks, or visit Firmulate to watch the live company.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work your business actually needs

For a pool service company, an AI that spots a problem but never closes the loop could leave a customer waiting or a valuable job unfinished. Firmulate’s results show why model choice deserves a practical trial: Kimi K3 nearly matched the leader, while the field still showed a gap between sound diagnosis and follow-through. Before putting AI to work in customer service, scheduling or sales, see how it handles your own difficult week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Managing Phosphates and Nitrates in Pool Water

Keeping phosphates and nitrates in check is essential for a clear pool; discover effective strategies to prevent algae growth and maintain pristine water.

Organic Load: Sunscreen, Leaves, and the Oxidation Budget

Diving into organic load from sunscreens and leaves reveals complex impacts on ecosystems, and understanding oxidation budgets is crucial for environmental health.

CYAnuric Acid (CYA): UV Shield Vs Chlorine Lock

Understanding whether CYA acts as a UV shield or causes chlorine lock can help you optimize pool maintenance—discover the key differences to keep your water crystal clear.

CO2 pH Control: When CO2 Systems Beat Acid Feeders (and When They Don’t)

Narrowing down when CO2 pH control surpasses acid feeders can save your system; discover the key factors that influence their effectiveness.