
Imagine testing a pool cleaner that claims to be the best, but it can’t even navigate a simple obstacle course. Now, replace that with AI managing a small company’s crises. The results might surprise you and highlight what really counts when trusting an AI — not just how well it chats, but whether it can finish the job honestly and thoroughly.
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the New Benchmark: More Than Just Scores
At first glance, AI models in recent experiments earned high scores — with the top reaching 95 out of 100. But the real story isn’t just about the numbers. Instead, it’s about what those numbers reveal when models face a simulated week filled with genuine business crises, temptations to cheat, and the need for integrity.
What the Numbers Say
The “Crucible League” results show a fascinating baseline: a do-nothing model scores 26 points. This might seem low, but it’s a crucial reference point. It means that even an AI that doesn’t actively work on solving problems can achieve partial credit just by doing the basics. More importantly, it underscores that partial progress is valued — every correct step counts — but a breach of trust caps the total score.
Why Partial Progress Matters
In the real world, no AI is perfect. So, the benchmark rewards models that can identify problems and handle them, even if not flawlessly. For instance, all tested models recognized every crisis and refused manipulative tricks. But only two managed to finalize and sign a significant deal, demonstrating that honesty and thoroughness matter more than just scoring high in isolated tasks.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deep Dive: The Hidden Weaknesses and Trust Issues
The most intriguing revelation was not in the obvious customer interactions but buried deep within internal files. Models that accessed these files successfully closed a €55,000 deal, adding over €4,500 in monthly recurring revenue. This shows that the true challenge isn’t just reacting to visible crises but understanding and trusting internal data that guides decision-making.
Social Engineering and Honesty Under Pressure
The experiment also included staged social engineering attacks, such as fake CEO messages and reporter tricks. All models refused to be manipulated, with Kimi K3 explicitly treating the requests as potential impersonation. This demonstrates a fundamental value: models that prioritize trustworthiness are crucial, especially in scenarios where deception could compromise decisions.
internal data security AI software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Tells Business Leaders
For those considering AI for customer service, support, or strategic planning, the key takeaway isn’t just whether the AI can generate persuasive language. It’s whether it can complete complex tasks honestly, read critical internal documents, and resist manipulation under pressure.
The Live Experiment in Action
Firmulate hosts a real-time, visible simulation of an AI managing a small operational company. With 13 synthetic employees, real revenue mechanics, and a cash burn of €105,000 per month, it’s a sandbox for testing management quality, not just chat skills. Every decision is versioned and auditable, providing transparency that’s often missing from other AI evaluations.
Performance Without Excessive Fine-Tuning
The models tested range from highly thorough participants, like Opus 4.8, to newcomers like K3. Even when running without effort parameters, models like K3 showed competitive discipline. This indicates that trustworthiness and integrity are less about tuning and more about inherent design qualities.
AI social engineering resistance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Thoughts: Trust Is the New Benchmark
The takeaway from these experiments is clear: in business, AI models must do more than just generate convincing text. They must finish what they start, diligently read internal data, and resist manipulation. Partial progress counts, but breaches of trust hit the score hard — and rightly so. As companies increasingly depend on AI for decision-making, these benchmarks serve as a reminder that honesty, thoroughness, and reliability are the true measures of value.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
