firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI that, instead of driving your business forward, simply does nothing — yet still scores 26 out of 100 in a rigorous industry benchmark. For interior designers or furniture retailers, this might sound like a nightmare scenario. But surprisingly, it’s a revealing reality about how we measure AI’s readiness to manage real-world workplace decisions.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Zero Isn’t Zero in AI Scores

When evaluating artificial intelligence in business settings, it’s tempting to think a ‘do-nothing’ approach would score a perfect zero. But in a recent public experiment conducted by Firmulate, a baseline model that deliberately refrained from acting scored 26 points. This isn’t a flaw; it’s a feature of how these benchmarks are built to reflect real-world complexities.

The Partial Progress Penalty and Trust Caps

The scoring system recognizes partial progress — even if the AI doesn’t complete the entire task, it can still earn points for correctly identifying crises or refusing manipulative tactics. However, the system also imposes a strict cap: if the AI breaches trust — say, by agreeing to a manipulation or signing a dubious deal — its total score is immediately capped at that breach point. This ensures honesty is always valued over short-term gains.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Benchmark Looks at Real Company Dynamics

The experiment involved four leading frontier models, each tasked with managing a small software company during its worst week. The same set of customers, crises, and temptations faced each model — an apples-to-apples test of decision-making quality. Every decision was recorded and auditable, mimicking real business processes.

Trust and Performance in Practice

Remarkably, all four models identified every crisis and refused every manipulation attempt, demonstrating a high level of integrity. Yet, only two managed to close a deal worth €55,000, earning the full score for that task. The other two, despite diagnosing correctly and pitching effectively, did not obtain signatures — resulting in a lower score.

The Hidden Weakness: Reading Critical Files

Deep within the data, a subtle but decisive weakness emerged: models that read and reference specific company files performed significantly better in closing deals. In fact, the models that examined two document references in the company’s files succeeded in securing the full deal value (+€4,583 MRR), while others missed this crucial detail.

Social Engineering and Ethical Decision-Making

The models faced staged social engineering attacks, including fake CEO messages escalating in severity and reporter tricks requiring a simple yes/no. All five models refused to be manipulated, trusting their protocols and reasoning: “Treat the request as a suspected approval-bypass / possible impersonation,” Kimi K3 explained. This shows that, in simulated environments, AI can uphold ethical guidelines under pressure.

The Live Business Simulation: Real Money, Real Risks

The experiment is not just theoretical. It plays out in a real-time, live simulation of a company with 13 synthetic employees, managing €2.3k in monthly recurring revenue against a burn rate of €105k. The platform, accessible at firmulate.com/live, allows users to watch AI decisions unfold daily, with over 680 self-learned rules and every workday versioned for analysis.

Insights from the Latest Participants

The most thorough model, Opus 4.8, with over 80 learned rules and deep analyses, ranked last in the scoring despite its sophistication. It left deals on the table and failed to escalate issues properly, revealing that even the most comprehensive AI can slip in discipline without precise guidance. Meanwhile, the K3 model, operating without an effort parameter, performed remarkably well, highlighting the importance of default fairness and trustworthiness.

Implications for Business Decision-Making

For interior designers, furniture retailers, or any business considering AI support tools, the key takeaway is clear: the real value isn’t just in how well an AI can generate text or ideas. It’s in whether it can finish what it starts, stay honest under pressure, and deeply understand your unique data — like a carefully curated client file or project blueprint.

As the experiments demonstrate, even a do-nothing baseline scores 26 points because the scoring rewards honest identification and refusal of manipulation, not just action. It’s a reminder that trustworthiness and thorough understanding are the true measures of AI readiness in business.

What Should You Expect From an AI Workforce?

Before deploying AI into your interior design or retail workflow, consider running a wargame — a simulated test to evaluate how the AI handles crises, manipulations, and trust issues. Firmulate offers such pilot experiments, allowing businesses to observe AI decision-making in a controlled, risk-free environment, ensuring they’re not just getting a clever chatbot but a reliable partner.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Remote Control of Appliances: Managing Devices From Anywhere

Beyond convenience, remote appliance control offers endless possibilities for smarter, more efficient home management—discover how to take advantage today.

Air Purifier Placement: The One Spot That Makes It Work Better

Great placement can maximize your air purifier’s effectiveness—discover the one spot that makes it work better and why it matters.

Build vs Buy a Prebuilt AI Workstation

Struggling to choose between building or buying your AI workstation? Discover the pros, cons, costs, and what suits your workload best in this detailed guide.

Smart Lighting Ideas for Modern Homes

Find out how smart lighting can transform your modern home into a vibrant oasis, creating the perfect atmosphere for every occasion. Discover the possibilities!