A standard for the consequence, not just the system

Most AI registers grade the system. Few grade what happens if it is badly wrong

A live register with a named owner tells you what is running and who answers for it. It does not, by itself, tell you which of those systems could plausibly contribute to serious injury or major loss if its judgement failed in an operation that touches physical assets, safety-critical processes or experienced workforces. We grade that separately, set the thresholds that define it, and build the incident response that follows once a system crosses them.

A threshold most registers do not have

In July 2026, Illinois became the first US state to legislate an independent audit requirement for frontier AI developers, through the Artificial Intelligence Safety Measures Act. Alongside New York and California, whose earlier laws already require published safety protocols, the Act defines catastrophic risk precisely: a foreseeable and material risk that a model's development, storage, use or deployment materially contributes to death or serious injury to more than fifty people, or more than a billion dollars in property damage, arising from a single incident. From January 2028, large frontier developers must retain an independent third party, one with genuine technical expertise in the field, to audit their compliance annually. A critical safety incident has to be reported within seventy-two hours as standard, tightened to twenty-four when the risk is imminent.

That law regulates the handful of organisations building frontier models, not the far larger population deploying them. But the mechanism it legislates, a written threshold for what counts as catastrophic, independent review of whether the organisation is managing it, and a timed obligation to report when something crosses the line, is not specific to frontier developers. It is good practice for any organisation whose AI estate reaches into work where a bad decision has a physical or financial consequence measured in more than an unhappy customer, and most of those organisations do not have it written down anywhere.

What we grade

For every system on the estate's live register, we ask a question the register's ordinary fields do not answer: if this system's judgement were badly wrong, on its worst plausible day, what could that contribute to.

ReachWhat the system can actually touch or influence
Whether the system's output is advisory, read by a person who decides, or whether it can act directly, adjusting a control, releasing a batch, scheduling a task, on a process where a wrong action has physical or safety consequences rather than only an inconvenient one.
SeverityA defined threshold, not a feeling
A written line, agreed with the operation rather than assumed by the AI team, for what would count as a serious injury, a major loss, or a materially unsafe outcome in this specific environment. The Illinois figures are one workable reference point; the right threshold for a given estate is set with the people who own the operational risk.
Independence of the checkWho confirms the grade is honest
Whether the person grading the system's risk has any stake in it being graded low. Where the answer is no, or where the workflow's own severity threshold is crossed, the grade is confirmed by someone outside the team that built or operates the system, in proportion to what is at stake rather than as a blanket audit requirement.

The response protocol

A system graded above the threshold gets three things a low-graded system does not need: a defined incident classification so a genuinely serious event is recognised as one rather than logged as an ordinary bug, a named escalation path with a person who can suspend the system without waiting for a committee, and a reporting timeline written down and rehearsed rather than improvised. The seventy-two-hour standard window and the twenty-four-hour window for imminent risk, drawn from the Illinois Act, are a sound default shape for that timeline even where no regulator requires it, because the discipline of having a clock already running is what keeps a live incident from losing its first, most useful hours to uncertainty about who decides.

None of this is paperwork produced after the fact. The classification and the escalation path are designed and agreed before the system is graded high enough to need them, tested against a rehearsed scenario the way a fire drill is tested against a scenario nobody hopes to use, and reviewed on the same cadence as the grade itself.

Who this is for

This standard earns its cost where AI already reaches into physical operations, safety-critical processes or decisions an experienced workforce currently makes by judgement: condition monitoring on live plant, scheduling or dispatch in an environment with real failure modes, or systems whose advice feeds a safety-relevant decision even where a human signs it off. An estate that is still confined to drafting, retrieval and internal knowledge work does not need this standard yet, and applying it there is effort spent ahead of its value, in the same way instrumentation built before there is volume to attribute is effort spent ahead of its value.

Where this connects

The grade itself is one column on the live register established as part of AI governance and ownership, and kept current under managed AI and technology. Where the concern is not what a system might do but whether the model behind it is telling the truth about what it is doing, the check that reaches inside the model itself is reading internal model state. And where the question is not severity but supply, whether the model this workflow depends on could simply stop being available, that is graded separately in grading model dependency risk. A system can clear all three checks and still be the wrong system for the job; it cannot be the right one without them.

Bring us the system nobody has asked 'what if it's badly wrong' about.

Tell us what it touches and who would be affected on its worst plausible day. We will show you how we would grade it and what a real response protocol would take to build.

Your message goes to the people who would do the work.