
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would an AI keep a pool business afloat through its worst week?
A pool company depends on more than clear water: customers need timely help, staff need direction and costly decisions have to hold up under pressure. For any business considering AI in its customer service or operations, polished answers are only part of the story. Firmulate’s live experiment puts models in charge of the same struggling software company and watches what they do when the week goes wrong.
A leaderboard with a narrow gap at the top
In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western models in the field. That is a close contest, not a settled hierarchy—and a reminder that choosing an AI model without testing it on your own work is a bet.
The company in the experiment faced €105,000 in monthly burn against €2,300 in monthly recurring revenue. Each frontier model faced the same customers, crises and temptations. Firmulate versions every workday’s decisions, making the experiment watchable as it unfolds. The company includes 13 synthetic employees, a public cash countdown and more than 680 self-learned playbook rules. The baseline score for doing nothing was 26; partial progress counted, but a single breach of trust capped the total.
Finding the clue was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between understanding the opportunity and completing the sale was the experiment’s sharpest finding: “Same diagnosis, same pitch — no signature.”
The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found that security-relevant detail, won the deal and saved the churning customer. It resisted all three baits and made just one deviation, the cleanest discipline in the field.
All five models also refused a staged social-engineering attempt: fake CEO messages escalated over three stages, followed by a reporter asking “just one yes/no, on background.” K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a pool or patio business, where an AI might handle customer details or service requests, that kind of restraint matters alongside speed and sales.
Thoroughness does not guarantee a finish
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models: good analysis did not always turn into a completed action.
There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate’s results are a useful prompt to test models against real business decisions, while keeping that difference in mind.
The experiment’s 242 real, unedited management decisions also power a “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Readers can follow the experiment and see the full findings at Firmulate’s benchmarks, or visit Firmulate to watch the live company.

Test the work your business actually needs
For a pool service company, an AI that spots a problem but never closes the loop could leave a customer waiting or a valuable job unfinished. Firmulate’s results show why model choice deserves a practical trial: Kimi K3 nearly matched the leader, while the field still showed a gap between sound diagnosis and follow-through. Before putting AI to work in customer service, scheduling or sales, see how it handles your own difficult week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
