AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would an AI keep a pool business afloat through its worst week?

A pool company depends on more than clear water: customers need timely help, staff need direction and costly decisions have to hold up under pressure. For any business considering AI in its customer service or operations, polished answers are only part of the story. Firmulate’s live experiment puts models in charge of the same struggling software company and watches what they do when the week goes wrong.

A leaderboard with a narrow gap at the top

In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer beat three of the four Western models in the field. That is a close contest, not a settled hierarchy—and a reminder that choosing an AI model without testing it on your own work is a bet.

The company in the experiment faced €105,000 in monthly burn against €2,300 in monthly recurring revenue. Each frontier model faced the same customers, crises and temptations. Firmulate versions every workday’s decisions, making the experiment watchable as it unfolds. The company includes 13 synthetic employees, a public cash countdown and more than 680 self-learned playbook rules. The baseline score for doing nothing was 26; partial progress counted, but a single breach of trust capped the total.

Finding the clue was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between understanding the opportunity and completing the sale was the experiment’s sharpest finding: “Same diagnosis, same pitch — no signature.”

The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found that security-relevant detail, won the deal and saved the churning customer. It resisted all three baits and made just one deviation, the cleanest discipline in the field.

All five models also refused a staged social-engineering attempt: fake CEO messages escalated over three stages, followed by a reporter asking “just one yes/no, on background.” K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a pool or patio business, where an AI might handle customer details or service requests, that kind of restraint matters alongside speed and sales.

Thoroughness does not guarantee a finish

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models: good analysis did not always turn into a completed action.

There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate’s results are a useful prompt to test models against real business decisions, while keeping that difference in mind.

The experiment’s 242 real, unedited management decisions also power a “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Readers can follow the experiment and see the full findings at Firmulate’s benchmarks, or visit Firmulate to watch the live company.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work your business actually needs

For a pool service company, an AI that spots a problem but never closes the loop could leave a customer waiting or a valuable job unfinished. Firmulate’s results show why model choice deserves a practical trial: Kimi K3 nearly matched the leader, while the field still showed a gap between sound diagnosis and follow-through. Before putting AI to work in customer service, scheduling or sales, see how it handles your own difficult week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Difference Between Clear Water and Truly Safe Water

Nakedly clear water may deceive, but discovering how to ensure true safety requires understanding hidden contaminants and proper testing methods.

The Metal Problem That Starts With Fill Water, Not Equipment

Metal ions in fill water silently threaten your system’s longevity, and understanding their impact is key to preventing costly damage—continue reading to learn more.

Breakpoint Chlorination: The Equation Behind “shocking”

AIThis post was created with the assistance of artificial intelligence (AI).Breakpoint chlorination…

Base Additions: Soda Ash Vs Baking Soda—When Each Applies

Gaining clarity on whether to use soda ash or baking soda can transform your household and industrial projects—discover which one fits your needs best.