firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI so indifferent that it barely tries — yet it still earns 26 out of 100 points on a demanding business benchmark. For beauty and personal care companies relying on AI for customer service, supply chain management, or marketing, understanding what this score really means is crucial. It’s not just about how well AIs chat or recommend products — it’s about whether they can truly finish what they start, stay honest under pressure, and deliver measurable results.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get beauty and skincare favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

At the heart of recent AI testing by the firmulate.com live experiment, four leading AI models faced the same intense scenario: managing a small software company’s worst week. Every decision was transparent, every crisis real, and every temptation to cheat was monitored. The goal? Measure their ability to handle real-world business pressures, not just generate convincing text.

One surprising result emerged: even the most indifferent AI, in a baseline run that did nothing, scored 26 out of 100. This ‘do-nothing’ score might seem trivial, but it reveals important truths about AI benchmarking. Since partial progress counts, the score isn’t zero, and it highlights that an AI’s ability to identify issues and act honestly isn’t solely about making progress — it’s about avoiding missteps. A single breach of trust caps the overall grade, setting an honest lowest bar for performance.

This baseline score underscores a key point: AI models are evaluated on their ability to detect and react to crises. In this experiment, all four models spotted every crisis and refused every manipulation attempt. For example, when fake CEO messages escalated over multiple stages, all models refused to proceed — Kimi K3 reasoned, “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a fundamental trust in AI decision-making under pressure.

Interestingly, the biggest competitive weakness was not in customer interactions but in internal document handling. The models that read two document references deep into the company’s files secured the deal at full price — worth over €4,583 MRR. This indicates that thorough understanding and access to internal data are vital for closing high-value deals. Conversely, superficial reading or failure to access critical information cost some models the opportunity.

Beyond decision-making, the experiment also tested social engineering scenarios, such as staged CEO messages and reporter tricks. All models refused to be manipulated, reinforcing their resilience against common tactics to bypass controls. The analysis aligns with Kimi K3’s reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Meanwhile, the live environment simulates a real company with 13 synthetic employees, burning €105k monthly against a revenue stream of just €2.3k. Every workday, the system learns and records decisions, making it a transparent, watchable test bed for AI performance. This setup illustrates that true readiness goes beyond chat quality; it’s about consistent, trustworthy work that withstands pressure.

Notably, the most thorough participant, Opus 4.8, with over 80 learned rules, finished last. It left a deal on the table and showed discipline slips, such as writing attempts into a locked department instead of escalating. This demonstrates that depth of analysis alone doesn’t guarantee high performance — discipline and process adherence matter just as much.

For business leaders, especially in beauty and personal care, the key takeaway isn’t how eloquently an AI can chat. It’s whether it can finish what it starts, read your files effectively, and stay honest when stakes are high. The benchmark’s lowest score of 26 reveals that baseline performance is surprisingly high — even minimal effort gets you somewhere, but only trust, discipline, and thoroughness lead to winning high-value deals.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

For beauty brands integrating AI, the real test isn’t just in shiny demos or smooth chats. It’s in how AI handles crises, reads internal data deeply, and remains honest under pressure. The firmulate.com benchmark shows even a do-nothing baseline scores above zero, emphasizing the importance of trust and discipline for AI to truly add value. As AI becomes more embedded in business operations, understanding these fundamentals will be key to choosing the right partner and securing measurable results.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Miu Miu Is Bringing Back An Archival Favorite

Miu Miu is launching a new collection that revisits a popular archival design, blending vintage elements with modern fashion. The move highlights the brand’s nostalgic approach.

It’s All About The Polo Sweater This Fall

This fall, fashion experts highlight the polo sweater as the must-have wardrobe staple, driven by rising search interest and ongoing style discussions.

WWD X FN X Beauty Inc Women In Power 2026

WWD, FN, and Beauty Inc will host the Women in Power 2026 event to celebrate influential women shaping the beauty industry, scheduled for 2026.

15 New Beauty Products Our Pros Are Obsessed With For September

Discover the top 15 beauty products experts are loving this September, highlighting emerging trends and favorites shaping the season’s skincare and makeup routines.