A method, not a hand-off
A written requirement is not what a veteran actually knows
The knowledge that decides whether an AI system is trustworthy rarely lives in a document. It lives in the exceptions a specialist has learned to catch, the judgement calls nobody wrote down because they seemed obvious at the time, and the near-misses from thirty years of product cycles that never made it into a process diagram. We elicit that knowledge directly from the people who hold it, and build the loop that keeps a system honest against it as the work changes.
The document is not the knowledge
Requirements documents describe the process as it is meant to run. Specialist judgement lives in the gap between that description and what actually happens: the case that looked routine but was not, the reading that contradicted the instrument, the supplier substitution that was fine this one time and not the next. None of that is negligence on the document’s part. A document cannot hold thirty years of exceptions, and it was never asked to.
Ford’s own experience makes the point concretely. The company has said it hired 350 veteran engineers over three years, brought back specifically to train younger staff and to improve AI tools that had not solved costly quality problems on their own. A Ford executive was direct about why: design requirements alone had not been enough, and experienced people had to train the tools using knowledge built up across many product cycles. That is a company discovering, at real cost, that the specification and the expertise are not the same artefact, and that only one of them can be handed over in a folder.
What elicitation looks like
Asking a specialist to describe their expertise produces a worse version of the same problem as the requirements document. People are good at their judgement and poor at narrating it in the abstract; asked what they know, they tend to state the rule they would defend rather than the exception they would actually apply. So elicitation does not start with a conversation about their process. It starts with their real cases.
We work alongside the specialist on live material: the report they are reviewing, the decision they are about to make, the output a system has already produced. We ask what they would change and why, case by case, and we pay closest attention to the corrections that surprise us, because those are the ones a generic description would never have surfaced. What accumulates is not a summary of their expertise. It is a record of the actual judgement calls, tied to the actual cases that produced them.
The correction loop
A single round of this is a good start and a poor foundation. Knowledge captured once goes stale the moment the operation changes, which is most of the time, so the elicitation has to keep running rather than conclude. We build it as a loop: the system produces output against real cases, the specialist reviews it, the gap between what they would have done and what the system did becomes the next round of correction, whether that lands as training material, a prompt, a rule the system checks against, or a case added to the set it is evaluated on.
- ProduceThe system attempts the real case
- Working material, not a hypothetical, so the specialist is reviewing something with consequences rather than grading a demonstration.
- ReviewThe specialist judges it as their own work
- The same standard they would apply to a junior colleague's draft: not merely right or wrong, but what a competent person in their seat would have caught.
- CorrectThe gap becomes material
- Every disagreement is recorded as a case, with the specialist's reasoning attached, and folded back into what the system is trained, prompted or checked against.
The loop is what makes this a method rather than a one-off project. It also has a practical consequence worth stating plainly: it never really finishes, because the operation keeps producing new exceptions for as long as it keeps running, and a system that stopped learning from correction the day it launched will start drifting from the judgement it was built on.
Why not just an interview
An interview asks someone to introspect on knowledge that mostly operates below the level they narrate. It produces a plausible, tidy account of their judgement that is measurably thinner than the judgement itself, because the exceptions that matter most are exactly the ones that feel too obvious, too rare, or too hard to explain to mention unprompted. Working from real cases sidesteps that problem: the specialist is not asked to describe their expertise, only to apply it, which is the thing they are already good at.
A document handover fails for a related reason. It transfers whatever the author thought to write down at the time, frozen at that moment, with no mechanism for the operation’s next exception to reach the system at all. The correction loop is slower to set up than either shortcut and produces something neither of them can: a system whose behaviour keeps answering to the person whose judgement it is meant to carry.
Who stays in the room
The specialist does not hand over their knowledge and leave. They stay the named authority behind what the system asserts, reviewed and credited case by case, which is the same principle behind how we design AI programmes for the specialists and veterans who use them. An expert who watches their own corrections change the system’s behaviour arrives at ownership of it. An expert asked to fill in a form once and never consulted again has no reason to trust what gets built from it, and every reason to be right about that.
Bring us one piece of expertise that matters.
Name the judgement call that experienced people make correctly and nobody has ever managed to write down. We will show you what the elicitation and correction loop would look like for it, and what the first few rounds would produce.
Your message goes to the people who would do the work.
