A standard for a decision usually left implicit
A model choice is a supply decision, whether or not anyone grades it
Choosing a model chooses more than capability and price. It chooses whose decisions, commercial or governmental, can interrupt the workflow that now depends on it. We grade that dependency for every model an estate actually relies on before an outage or a policy change forces the question, and we design the routing that turns a bad grade into a manageable one.
The supply question nobody grades
Most model evaluations stop at three questions: what can it do, what does it cost, what is it allowed to touch. Those are the right questions for choosing a model on an ordinary day, and we ask them as a matter of course. None of them answers a different question that only surfaces on the day it stops being hypothetical: what happens to the workflow if this specific model is not there tomorrow, for reasons that have nothing to do with how well it performs. A model can be the best-evaluated choice available and still be a single point of failure nobody wrote down.
That gap is not exotic. A frontier model went dark for eighteen days in June 2026 because of an export control order, not an outage. It is the kind of interruption that a capability comparison will never catch, because capability was never the thing that failed.
What we grade
For every model a workflow that matters actually depends on, we grade three things.
- ConcentrationIs there really only one way to do this
- Whether a workflow that matters is served by exactly one model from exactly one provider, with no alternative that has ever been exercised on real cases rather than assumed to work.
- Provenance and jurisdictionWhose decision could interrupt this
- Whose commercial terms, hosting schedule or national export and distribution policy currently allow you to use this model on the terms you rely on today, and how much standing you have in the room where that could change.
- SubstitutabilityWhat has quietly grown around this specific model
- How much of the surrounding system, prompts tuned to its habits, an evaluation suite calibrated against its behaviour, routing thresholds set when it was the only candidate, would need to be rebuilt rather than reconfigured if it had to be replaced.
None of these is a reason to avoid a model. They are the questions that turn a dependency an organisation did not choose to notice into one it decided to accept, on terms it wrote down.
The grade
Grading ends in one of three states, recorded against the workflow rather than left as a general impression of the provider.
- Sole-sourced
- No alternative has been exercised for this workflow. This is sometimes the right position for low-stakes work, and it should never be the position an important workflow ends up in by accident. A sole-sourced grade requires a documented fallback plan and a date to revisit it, even when the honest answer for now is to accept the risk.
- Hedged
- A declared alternative exists and has been run against the same ground truth recently enough that its results are known rather than assumed, even if it costs more, runs slower or covers less of the task than the primary choice.
- Diversified
- The workflow already runs across more than one model or provider by design, worker and adviser, panel and judge, or a mixed private and frontier architecture, so no single supplier's decision can interrupt it outright.
Routing does the hedging
A grade is only useful if something can act on it, and the machinery that acts on it already exists for other reasons. Explicit routing between models is normally built for capability, governance and unit economics. The same routing, once it exists, is what turns a sole-sourced workflow into a hedged one: a declared alternative only counts if the system already knows how to call it, and an alternative that has never been exercised in production is a plan, not a hedge. Where a workflow's grade genuinely does not justify that machinery, the fallback can be simpler: a documented manual process and a named person who owns invoking it, reviewed on the same schedule as the grade itself.
Why this cuts both ways
The interruption above ran in one direction: a government restricting who could reach a model that was already available. The same risk can run the other way. Reuters reported in July 2026 that China's Ministry of Commerce had held exploratory discussions with Alibaba, ByteDance and Z.AI about restricting overseas access to advanced Chinese AI models, including a proposed tiered regime that would put basic open tools under a simple filing, more advanced technology under security review, and frontier models under domestic-only use. The reporting is consistent that nothing has been decided, that the scope remains under discussion, and that the options sketched apply chiefly to future releases rather than withdrawing models already distributed. Separately, a public legal dialogue inside China has been debating adjacent questions, including whether open source is straightforwardly pro-competitive, and some commentators argue that dialogue should not be read as a preview of settled policy. We take no position on whether these restrictions should happen or how the debate should resolve. The operational point holds regardless: an organisation that has built part of its estate on an open-weight model, expecting the supplier's next generation to keep arriving on the same terms as the last one, is depending on a policy decision made in a jurisdiction where it has no standing, in exactly the shape this standard exists to grade.
One distinction matters for the grade itself. Weights already downloaded and deployed inside your own boundary keep working regardless of a later export decision; nothing reported would reach back and switch off a model you already run. What a restriction of this kind would remove is the next generation, not the last one. That is a renewal risk rather than an availability risk, and it belongs on the same register under provenance and jurisdiction: not "does this model work today" but "can I still expect the improved version I was planning to move to."
Where this connects
Grading a workflow's dependency is one input to the live register described in AI governance and ownership, kept current as part of managed AI and technology. The routing that acts on a grade is the same machinery covered in model orchestration. And the underlying argument, that capability you did not build yourself runs on terms you did not write, is developed further in the model you depend on is rented. Where a model's judgement, if badly wrong, could plausibly contribute to serious harm or major loss rather than only lost supply, that severity is graded separately in grading catastrophic AI risk.
Bring us the workflow you have never had to fail over.
Tell us which model it depends on and what happens if that model is not there one Monday. We will show you how we would grade it and what building a real alternative would take.
Your message goes to the people who would do the work.
