A method for turning elicited judgement into a specialised model
Sometimes the cheapest reliable model for a task is the one you train, not the one you rent
Eliciting a specialist's judgement produces a record of real cases and the corrections an expert made to them. One of the things that record can become is training material: a smaller model built to do one task the way your specialists do it, at a fraction of the cost of routing every case to a general model capable of everything. We treat post-training as its own engineering decision, made on evidence, not as the default next step after elicitation or the last resort when a general model is too expensive.
General intelligence is not always the right tool
The convenient default for a new task is a frontier model: broad, capable, available immediately, priced by the call. For a great many tasks that convenience is also the right engineering answer, and choosing it deliberately beats avoiding it out of principle. But a frontier model is solving a much broader problem than the one in front of it, and for a task that repeats at volume, in a narrow shape, that breadth is mostly unused capability the task is still paying for.
Eliciting a specialist's judgement already produces the raw material for a narrower answer. The correction loop generates real cases, the specialist's actual decision on each one, and the reasoning behind it, and that record can be turned into a prompt, a rule the system checks against, or training material for a model built specifically for the task. Post-training is what happens when the third option is the right one: a smaller model, trained on that record, doing one job as well as a much larger one, for a small fraction of the cost.
This is not a hypothetical trade. Bridgewater worked with Thinking Machines to fine-tune a model on its own analysts' financial judgement using Thinking Machines' Tinker platform; the resulting model reportedly reached roughly 85% average accuracy across the firm's core tasks at single-digit-dollar cost per task, against 74 to 78% for general frontier models costing roughly ten to ninety times as much on the same tasks, according to Thinking Machines' own published account of the project. Microsoft has separately reported an MAI model tuned for Excel-specific tasks matching a general frontier model's quality at roughly a tenth of the cost, and a model tuned on McKinsey's own material beating a general model on quality at a similar saving, according to Microsoft's own announcement. Both are the vendor's own reported results on the vendor's own benchmark, not independently audited figures, and we cite them as evidence that the pattern works at all, not as a guarantee of any particular saving.
What post-training needs
Post-training a model well is not a shortcut around the elicitation work; it depends on it having already been done properly.
- Enough graded casesA pattern, not a handful of anecdotes
- A body of elicited corrections large enough that what the model learns is the shape of the specialist's judgement rather than a few memorable exceptions. There is no fixed number; the test is whether held-out cases the model has not trained on are handled the way the specialist would handle them.
- A held-out evaluation setKept separate from training from the start
- Real cases the specialist has graded but that never enter training, so the model's performance can be measured honestly rather than checked against the material it learned from.
- A base model worth specialisingUsually open-weight, licensed for the purpose
- Post-training starts from an existing model rather than nothing, most often an open-weight model deployable inside your own boundary, chosen and licensed the same way any model in the estate is chosen: on capability, governance and unit economics.
The decision to specialise
Post-training earns its cost on some tasks and wastes it on others, and the difference is usually visible in advance.
- Specialise when
- The task is narrow and repeats at real volume; the cost, latency or availability of a general model is the binding constraint on the workflow; the judgement being encoded is stable enough to be worth freezing into weights for a while; and evidence from a real evaluation shows the specialised model actually beats the general model doing the same task, not merely that it is cheaper.
- Keep renting when
- The task's shape keeps changing faster than a model could be usefully retrained; volume is too low to justify the evaluation and maintenance a specialised model needs; or the judgement it would encode is still forming, in which case freezing it into weights would freeze in today's best guess rather than today's best practice.
Keeping it honest
A specialised model can drift from the judgement it was built on in exactly the way a prompted one can, and for the same reason: the operation it serves keeps producing new exceptions after the model stops changing. The same correction loop described in eliciting tacit expertise has to keep running against the trained model's real output, not only against the prompted system it may have replaced, and the held-out evaluation set needs refreshing as the operation changes, on the same discipline set out in measurement and instrumentation. A specialised model that stopped being checked the day it shipped is a snapshot of a specialist's judgement on the day it was trained, quietly going stale.
What this does not replace
Post-training is a third option alongside the two covered in model orchestration: choosing the right existing model for a task, and combining several existing models within a task. A specialised model can be the worker in a worker-adviser pairing, or the whole answer for a narrow task on its own; either way, it is still chosen on capability, governance and unit economics like anything else in the estate, and where it is your only capable model for a task, it is also the kind of dependency covered in grading model dependency risk. Owning the weights removes some supply risk and does not remove all of it: a specialised model still depends on the base model it was built from, and on the team that keeps retraining it as the operation moves on.
Bring us the task your specialists could grade in their sleep.
Name the judgement call your best people make correctly, over and over, at real volume. We will show you what turning it into a specialised model would take, and whether it is worth doing.
Your message goes to the people who would do the work.
