AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a pool cleaner that claims to be the best, but it can’t even navigate a simple obstacle course. Now, replace that with AI managing a small company’s crises. The results might surprise you and highlight what really counts when trusting an AI — not just how well it chats, but whether it can finish the job honestly and thoroughly.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the New Benchmark: More Than Just Scores

At first glance, AI models in recent experiments earned high scores — with the top reaching 95 out of 100. But the real story isn’t just about the numbers. Instead, it’s about what those numbers reveal when models face a simulated week filled with genuine business crises, temptations to cheat, and the need for integrity.

What the Numbers Say

The “Crucible League” results show a fascinating baseline: a do-nothing model scores 26 points. This might seem low, but it’s a crucial reference point. It means that even an AI that doesn’t actively work on solving problems can achieve partial credit just by doing the basics. More importantly, it underscores that partial progress is valued — every correct step counts — but a breach of trust caps the total score.

Why Partial Progress Matters

In the real world, no AI is perfect. So, the benchmark rewards models that can identify problems and handle them, even if not flawlessly. For instance, all tested models recognized every crisis and refused manipulative tricks. But only two managed to finalize and sign a significant deal, demonstrating that honesty and thoroughness matter more than just scoring high in isolated tasks.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Deep Dive: The Hidden Weaknesses and Trust Issues

The most intriguing revelation was not in the obvious customer interactions but buried deep within internal files. Models that accessed these files successfully closed a €55,000 deal, adding over €4,500 in monthly recurring revenue. This shows that the true challenge isn’t just reacting to visible crises but understanding and trusting internal data that guides decision-making.

Social Engineering and Honesty Under Pressure

The experiment also included staged social engineering attacks, such as fake CEO messages and reporter tricks. All models refused to be manipulated, with Kimi K3 explicitly treating the requests as potential impersonation. This demonstrates a fundamental value: models that prioritize trustworthiness are crucial, especially in scenarios where deception could compromise decisions.

Amazon

internal data security AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Tells Business Leaders

For those considering AI for customer service, support, or strategic planning, the key takeaway isn’t just whether the AI can generate persuasive language. It’s whether it can complete complex tasks honestly, read critical internal documents, and resist manipulation under pressure.

The Live Experiment in Action

Firmulate hosts a real-time, visible simulation of an AI managing a small operational company. With 13 synthetic employees, real revenue mechanics, and a cash burn of €105,000 per month, it’s a sandbox for testing management quality, not just chat skills. Every decision is versioned and auditable, providing transparency that’s often missing from other AI evaluations.

Performance Without Excessive Fine-Tuning

The models tested range from highly thorough participants, like Opus 4.8, to newcomers like K3. Even when running without effort parameters, models like K3 showed competitive discipline. This indicates that trustworthiness and integrity are less about tuning and more about inherent design qualities.

Amazon

AI social engineering resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Thoughts: Trust Is the New Benchmark

The takeaway from these experiments is clear: in business, AI models must do more than just generate convincing text. They must finish what they start, diligently read internal data, and resist manipulation. Partial progress counts, but breaches of trust hit the score hard — and rightly so. As companies increasingly depend on AI for decision-making, these benchmarks serve as a reminder that honesty, thoroughness, and reliability are the true measures of value.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Flocculants Vs Clarifiers: Particle Science for Clear Water

Learning the differences between flocculants and clarifiers reveals how particle science improves water clarity—discover which method suits your needs.

Why Temperature Swings Can Throw Water Balance Off Faster Than Expected

Narrow temperature swings can disrupt your body’s water balance quickly, and understanding why is key to staying properly hydrated in changing conditions.

UV Sanitizers: What They Kill, What They Don’t, and Why Chlorine Still Matters

Learn why UV sanitizers can’t replace chlorine for complete disinfection and how combining both methods ensures safer, more effective hygiene.

Comparative Costs of Different Sanitizing Systems

Investigating the costs of sanitizing systems reveals key differences that can impact your choice—continue reading to find out which option is best for you.