A standard for the moment before trust is earned

A new model release earns trust before it earns a role

Several genuinely significant models now arrive most weeks, each announced with claims that outrun what has actually been tested. Before any of that capability reaches a client's estate, we separate what a release has demonstrated from what it has only been said to do, then find out which interaction style, cost profile and work role it genuinely fits, tested against our own trusted cases rather than the vendor's own benchmark.

Why triage replaced adoption

A few years ago, a new frontier model was an event: rare enough that evaluating it properly was a reasonable use of a week. That pace no longer exists. Multiple releases a lab would once have called significant now arrive in an ordinary week, from more than one lab at once, each with its own benchmark chart and launch post. Adoption by hype, trying whatever shipped most recently because it shipped most recently, stops being a minor inefficiency at that cadence and becomes a way of running an estate on rumour.

The question a serious buyer needs answered has changed shape with it. It is no longer just "which model is smartest", a question that a single leaderboard number implies has one answer. It is which interaction style, cost profile and work role each new release actually fits, for the specific workflows an organisation runs today. Two releases can both be genuinely strong and still deserve completely different jobs.

A claim is not a capability

The industry around model releases runs on three different kinds of statement, and treating them as interchangeable is the first mistake worth ruling out. A lab's stated outlook about what future capability will look like, framing sometimes summarised as there being "no ceiling" on what is coming, is a forecast about direction, not a demonstrated result about the model in front of you. A rumour about a forthcoming model's scale, timing or capability is a report, sourced or otherwise, until it is independently verified against something a third party can actually check. Neither is evidence that a specific release, available today, performs a specific task reliably.

We are not naming a particular rumoured model here, because the point is not any one of them. It is the general failure mode: treating a lab's stated ambition, or a report that has not yet been confirmed, as though it were a tested fact about what an organisation can rely on this quarter. A release deserves a role once it has been tested, not once it has been announced.

What we test, and against what

Once a release exists to be tested, the method is the same every time, and it is deliberately narrower than a vendor's benchmark suite.

Our own ground truth, not the vendor's benchmarkTested on the organisation's trusted cases
A vendor's benchmark is built to make the vendor's model look good on tasks the vendor chose. We run a new release against cases the organisation already trusts the answer to, graded by the same standard as every other model in the estate, before its result means anything to that organisation's decisions.
Interaction style and role, not a single rankingFit over a crown
We identify which role a model actually suits: a background implementer working under supervision, a model in a supervising or advisory position over other work, or a conversational or voice layer users talk to directly. A model can be an excellent implementer and a poor adviser, or the reverse. Crowning one model best in the abstract answers a question nobody needs answered.
Cost and reliability, treated as first-classNot an afterthought to capability
Cost per task and the rate at which a model fails, including in ways a raw capability score never sees, are measured alongside accuracy from the start, not added later as a footnote once a model has already been chosen on capability alone.

The fine print behind a headline number

An independent evaluation from Artificial Analysis, running its AutomationBench-AA benchmark built on Zapier's AutomationBench simulated office-tool task set, covering tools such as Gmail, Sheets and Slack, illustrates why a single headline number is not the same as a trustworthy answer for a given role. Grok 4.5 was independently reported completing these simulated workflows at around a 51 per cent success rate, at roughly $0.34 per task. Comparator frontier models in the same evaluation completed around 48 per cent of tasks, at roughly $1.35 to $1.46 per task. On capability and cost alone, that reads as a clear result: similar success rate, a fraction of the price.

The same evaluation also found that Grok 4.5 had the highest rate of guardrail and business-rule violations among the models compared. For a task where the rule being broken does not matter much, that finding is close to irrelevant and the cost advantage stands on its own. For a task where the rule exists precisely because breaking it is expensive, sending an email that should not have gone out, editing a cell a process depends on, posting to a channel that should have stayed private, the same model becomes the wrong choice at any price, regardless of how it ranks on success rate. This is the distinction our method is built to catch before a workflow depends on it, and it is exactly why cost and reliability are tested as first-class rather than checked afterwards.

Where this connects

This standard covers the earlier moment: deciding whether a brand-new release deserves trust at all. Two related standards pick up once that question is settled. The ongoing, per-task choice among models an estate already trusts is the right model for each job, covered in technology and product delivery. And once a model is relied upon for a workflow that matters, the separate question of what that reliance exposes an organisation to, which supplier, which jurisdiction, what happens if the supply is interrupted, is covered in grading model dependency risk. This page exists before either of those: a release has to earn trust before it can be routed to a job, and long before an estate can depend on it. All three test what a model does. Where the model can be instrumented, a further and separate check on what it privately represents while it works is reading internal model state.

Bring us the release everyone is talking about.

Tell us what it is being asked to do and what you already trust the answer to. We will show you how we would test it, and what we would tell you before you let it near real work.

Your message goes to the people who would do the work.