A standard for judging AI output
The bar a deliverable has to clear, not the task it completed
A piece of AI-produced work can finish a task and still not be good enough to use. We grade individual deliverables, the report, the code, the analysis, the design, against a specific standard: would this be good enough that a real paying client, in that profession, would accept it? Not whether the steps ran. Not whether the answer looks plausible. Whether it clears the bar the market already sets.
Finishing is not the same as being good enough
Most evaluation of AI output quietly measures the wrong thing. Did it produce an answer. Did the answer look complete. Did the process run without erroring. Those are useful engineering checks, and none of them tell you whether the thing produced is fit to put in front of somebody who is paying for the result and has their own professional standard for what good looks like.
A report can be structurally complete and still miss the finding a competent analyst would have caught. A piece of code can run and pass its own tests and still be something no engineer would sign off. Task completion asks whether the machine did something. The professional acceptance bar asks a harder and more useful question: whether a person who does this for a living, and whose reputation rides on the answer, would hand this exact piece of work to their client.
The bar, defined
We define the standard plainly, because a standard that cannot be stated plainly cannot be applied consistently. A deliverable clears the professional acceptance bar when someone qualified to judge that kind of work, working blind and without knowing which output came from which source, would judge it fit to deliver to a paying client in that profession, at that client’s expense, under that client’s scrutiny.
That is a binary test, not a score. A deliverable either clears the bar or it does not, because a client either accepts the work or sends it back. Grading on a sliding scale of helpfulness or fluency measures something more forgiving than the market actually applies, and it is the market’s standard we are interested in, not a more comfortable one of our own.
How we apply it
The method is a paired comparison, not an isolated review. A genuine professional’s own deliverable for the same brief becomes the reference point, the thing a client already accepted and paid for. The AI-produced output is judged against that same brief, by someone qualified to make the judgement a client in that field would make, without being told which piece came from which source.
- The referenceWhat good already looks like
- A real professional's own deliverable for the same brief, the standard a client has already accepted and paid for, not a description of what one might look like.
- The judgeQualified to hold the standard
- Someone who can tell competent work from adequate-looking work in that specific discipline, evaluating blind so the source of the output cannot colour the verdict.
- The verdictPass or fail against defined criteria
- A binary decision built from criteria fixed before the review, not a fluency score. It either clears the bar a paying client would apply, or it does not.
Fixing the criteria before review matters as much here as it does anywhere else we apply evaluation. A judge who works out what good means while looking at the answer will find reasons the answer they liked was good, which is exactly the bias a blind, criteria-first review exists to remove.
Why this matters now
The reason this standard needs stating explicitly is that capability and professional readiness are moving at different speeds, and the gap between them is where the risk sits. The Center for AI Safety’s Remote Labor Index, which tests models against real, economically valuable freelance work by comparing their output to a paid professional’s gold-standard deliverable, reported that frontier performance more than quadrupled in under eight months.
The same benchmark reported the leading model clearing that professional bar on a little over one in six of the tasks tested. Read generously, that is remarkable progress from a standing start. Read carefully, it means roughly five out of six pieces of professional-grade work still failed a bar that a real freelancer clears often enough to get paid. Both readings are true at once, and a business that only hears the trend line and not the failure rate is the business that ships work its own client rejects.
How it relates to our other standards
This is a sibling standard to two we already apply, not a replacement for either, because each answers a different question. The Evidence Sprint grades a workflow: does AI assistance change the outcome of a process your operation runs, measured against thresholds you agreed before the trial. That is a question about effect, tested case by case across a whole workflow.
The professional acceptance bar grades a single piece of output on its own terms: is this particular deliverable good enough to hand over. A workflow can show a real, positive effect on average while individual outputs still fail this bar often enough to need a human check before anything leaves the building, which is precisely the distinction a business relying on the average would miss.
It is a different exercise again from the evidence grading we apply to claims about a company or a portfolio, which asks how strongly a statement is supported (measured, verified, demonstrated, asserted) rather than whether a finished piece of work meets a professional standard. Claims are graded for how well they are supported. Deliverables are graded for whether they are good enough to use.
Where it gets used
This standard earns its place inside work we already do rather than existing as a separate offer. It is one of the checks that makes evaluation that catches quiet failure concrete for systems that produce discrete deliverables: reports, code, analyses, drafted decisions. It is also available as a grading method inside an Evidence Sprint when the workflow in question turns on the quality of a specific output rather than only on time or cost, so the sprint can tell you both whether the workflow got faster and whether what it produced would actually survive a client’s desk.
Bring us a deliverable that matters.
Tell us what the AI is producing, and who it ultimately has to satisfy. We will show you what grading it against a professional acceptance bar would involve, and what you would learn from the first review.
Your message goes to the people who would do the work.
