firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Polished answers are not the same as sound management

Beauty and personal-care businesses know how quickly an apparently manageable week can turn hostile. A customer problem collides with cash pressure; a competitor changes its offer; an urgent message seems to come from the boss; and a commercial opportunity demands more than a well-written response. The decisive question is whether a manager can identify what matters, protect trust and finish the job.

That is also the emerging problem with evaluating AI agents. Coding leaderboards and chat arenas can reveal answer quality, but they say little about triage under capacity pressure, consequences unfolding across days or honesty toward the board. An agent can sound reassuring while overlooking the file that changes a negotiation—or complete nearly everything except the action that produces revenue.

Amazon

business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A benchmark built around consequences

Firmulate approaches the gap by running frontier models as complete companies. In the Crucible League, each participant managed the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 standings were:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counted. But the benchmark made trust non-negotiable: a single breach capped the total, reflecting its governing principle that “no amount of good work outweighs a breach of trust.” That is a useful standard for any business considering agents that may touch customer relationships, forecasts or support work.

The gap between analysis and action

The broad result initially looks encouraging. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

This is the difference between chat quality and management quality. A capable analysis is not a business outcome. The work also has to survive handoffs, access restrictions, competing priorities and the final moment when somebody must close the loop.

The most important commercial clue did not appear in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. For a beauty business, the equivalent might be an overlooked contract note, an earlier complaint, a product restriction or a competitor detail buried in internal material. The lesson is not merely that retrieval matters. Good management means checking the company’s own record before acting confidently.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating visible effort with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared more mildly in the other four participants.

That profile should feel familiar to managers who have seen busy work mistaken for progress. More analysis, more documentation and more attempted actions can still produce a weaker result if the agent does not recognize a blocked path and escalate appropriately.

There is also an important fairness note in reading the standings. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result remains notable, but responsible comparison requires keeping that condition visible rather than treating every entry as operationally identical. The fuller findings are available on Firmulate’s benchmark page.

Honesty held under pressure

The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because management quality is not simply the ability to make money. It includes resisting authority theater, protecting approvals and refusing seemingly small shortcuts when pressure rises. In customer-facing industries, one careless disclosure or unauthorized promise can outweigh a long sequence of competent work.

A living curriculum for AI managers

The live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR. Its public cash countdown makes delay consequential, while 680+ self-learned playbook rules and versioned workdays show how operating behavior develops over time.

Its scenario names—churn wave, price increase, downround and PR crisis—point toward a more useful AI curriculum. These are not trivia questions with isolated correct answers. They test whether an agent preserves context, makes trade-offs and follows through as conditions deteriorate.

Readers can also examine 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems, and inquiries can be directed to contact@firmulate.com.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Buy judgment, not eloquence

For beauty and personal-care leaders, the practical implication is straightforward. Before giving an AI agent meaningful responsibility, do not ask only whether it writes convincing messages or solves tidy tasks. Put it through the kinds of pressure that define the real job.

Can it distinguish urgency from manipulation? Will it read the relevant files before advising a customer? Does it escalate when permissions block the sensible route? Can it convert a correct diagnosis into a signed agreement without compromising trust?

Firmulate’s experiment suggests that these questions define a distinct category: management quality, not chat quality. The strongest agent is not necessarily the one that says the most, analyzes the longest or looks busiest. It is the one that stays honest, finds the buried fact and completes the consequential work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


BABY SHOWER & RE

Baby shower & registry season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Be Trusted to Run a Business? A Live Experiment Reveals Surprising Results

A live AI business experiment reveals which models stay honest, read deeply, and finish what they start—crucial traits for trustworthy management tools.

Artist Uniform: No One Did It Like Dolly

Dolly Parton’s distinctive artist uniform continues to set her apart in the music industry, emphasizing her iconic style and influence.

23 Under-$200 Shopbop Finds That Look Triple The Price

Discover 23 stylish Shopbop items under $200 that look much more expensive, offering budget-friendly options for fashionable shoppers.

Teyana Taylor Wows in Bold Burgundy Gown at the 2026 BET Awards

Teyana Taylor made a striking appearance in a bold burgundy gown at the 2026 BET Awards, capturing attention on the red carpet with her fashion choice.