
Imagine hiring an AI that, instead of driving your business forward, simply does nothing — yet still scores 26 out of 100 in a rigorous industry benchmark. For interior designers or furniture retailers, this might sound like a nightmare scenario. But surprisingly, it’s a revealing reality about how we measure AI’s readiness to manage real-world workplace decisions.
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline: Why Zero Isn’t Zero in AI Scores
When evaluating artificial intelligence in business settings, it’s tempting to think a ‘do-nothing’ approach would score a perfect zero. But in a recent public experiment conducted by Firmulate, a baseline model that deliberately refrained from acting scored 26 points. This isn’t a flaw; it’s a feature of how these benchmarks are built to reflect real-world complexities.
The Partial Progress Penalty and Trust Caps
The scoring system recognizes partial progress — even if the AI doesn’t complete the entire task, it can still earn points for correctly identifying crises or refusing manipulative tactics. However, the system also imposes a strict cap: if the AI breaches trust — say, by agreeing to a manipulation or signing a dubious deal — its total score is immediately capped at that breach point. This ensures honesty is always valued over short-term gains.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Benchmark Looks at Real Company Dynamics
The experiment involved four leading frontier models, each tasked with managing a small software company during its worst week. The same set of customers, crises, and temptations faced each model — an apples-to-apples test of decision-making quality. Every decision was recorded and auditable, mimicking real business processes.
Trust and Performance in Practice
Remarkably, all four models identified every crisis and refused every manipulation attempt, demonstrating a high level of integrity. Yet, only two managed to close a deal worth €55,000, earning the full score for that task. The other two, despite diagnosing correctly and pitching effectively, did not obtain signatures — resulting in a lower score.
The Hidden Weakness: Reading Critical Files
Deep within the data, a subtle but decisive weakness emerged: models that read and reference specific company files performed significantly better in closing deals. In fact, the models that examined two document references in the company’s files succeeded in securing the full deal value (+€4,583 MRR), while others missed this crucial detail.
Social Engineering and Ethical Decision-Making
The models faced staged social engineering attacks, including fake CEO messages escalating in severity and reporter tricks requiring a simple yes/no. All five models refused to be manipulated, trusting their protocols and reasoning: “Treat the request as a suspected approval-bypass / possible impersonation,” Kimi K3 explained. This shows that, in simulated environments, AI can uphold ethical guidelines under pressure.
The Live Business Simulation: Real Money, Real Risks
The experiment is not just theoretical. It plays out in a real-time, live simulation of a company with 13 synthetic employees, managing €2.3k in monthly recurring revenue against a burn rate of €105k. The platform, accessible at firmulate.com/live, allows users to watch AI decisions unfold daily, with over 680 self-learned rules and every workday versioned for analysis.
Insights from the Latest Participants
The most thorough model, Opus 4.8, with over 80 learned rules and deep analyses, ranked last in the scoring despite its sophistication. It left deals on the table and failed to escalate issues properly, revealing that even the most comprehensive AI can slip in discipline without precise guidance. Meanwhile, the K3 model, operating without an effort parameter, performed remarkably well, highlighting the importance of default fairness and trustworthiness.
Implications for Business Decision-Making
For interior designers, furniture retailers, or any business considering AI support tools, the key takeaway is clear: the real value isn’t just in how well an AI can generate text or ideas. It’s in whether it can finish what it starts, stay honest under pressure, and deeply understand your unique data — like a carefully curated client file or project blueprint.
As the experiments demonstrate, even a do-nothing baseline scores 26 points because the scoring rewards honest identification and refusal of manipulation, not just action. It’s a reminder that trustworthiness and thorough understanding are the true measures of AI readiness in business.
What Should You Expect From an AI Workforce?
Before deploying AI into your interior design or retail workflow, consider running a wargame — a simulated test to evaluate how the AI handles crises, manipulations, and trust issues. Firmulate offers such pilot experiments, allowing businesses to observe AI decision-making in a controlled, risk-free environment, ensuring they’re not just getting a clever chatbot but a reliable partner.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
