firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Would you trust an AI with the family calendar?

Imagine handing an assistant access to school schedules, family finances and a message that appears to come from your partner. A polished answer would be reassuring. But the real test is whether it spots trouble, checks the details, resists pressure and follows through. That is the question behind Firmulate, a live experiment that puts AI models in charge of a software company during its worst week.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate gave each frontier model the same small software company, the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown and versioned workdays make the experiment watchable at Firmulate.

The final Crucible league, dated July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. The principle is plain: “no amount of good work outweighs a breach of trust.”

Amazon

enterprise AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Seeing the problem isn’t the same as finishing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between identifying the right answer and completing the work is the experiment’s central finding.

The decisive competitor weakness was tucked two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found the buried security fact, won the deal and saved the churning customer. It resisted all three baits and made one deviation, the cleanest discipline in the field. It beat three of the four Western frontier models in the final table.

The pressure included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful instinct wherever an AI might encounter a message asking it to skip a safeguard.

Opus 4.8 offers a counterpoint to the idea that more activity means better management. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four.

Amazon

AI ethical decision models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the result reaches beyond business

Parents already weigh similar questions when choosing tools for children and households: Does it protect private information? Does it check before acting? Does it keep its head when someone applies pressure? A company simulation cannot settle those questions for every family. But it makes one point visible: strong answers in a chat are not proof that an assistant will act reliably across a chain of decisions.

Firmulate says the live company has accumulated more than 680 self-learned playbook rules, and that every workday is versioned. Its benchmark page lays out the results and findings in plain language: see the Firmulate benchmarks. The experiment also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice.

Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

business AI simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The lesson for choosing an AI

Kimi K3’s second-place finish makes the field look open, while the unsigned deal shows why a leaderboard alone cannot answer every practical question. If an AI may touch a family calendar, a support queue or business records, test it on the kinds of decisions it will actually face. Choosing without your own test is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Premium Categories on Amazon Where “Deals” Can Mislead You

Never assume all Amazon “deals” are genuine; discover how to spot real savings and avoid marketing tricks.

Amazon Haul Finds Under $10 That Feel Like a Steal

Curious about amazing Amazon finds under $10 that feel like a steal? Discover budget-friendly treasures you won’t want to miss.

How to Shop Amazon More Intentionally Without Missing Good Finds

Discover how to shop Amazon more intentionally and find hidden gems without missing out on great deals or sacrificing quality.

AI’s Diligence Is Not Enough: Why Focus Matters More Than Volume in Business Decisions

Deep AI experiments reveal that thoroughness alone doesn’t ensure success. Prioritization and context-reading are key—less volume, more focus in family and business.