Platforms and tools

Sometimes the right model for the job is more than one model

Choosing a model per task answers most questions well. Some tasks are better served by combining models within the task itself: a cheaper model doing the work with a stronger one reviewing it, or several models attempting an answer that a separate process judges and combines. We build and govern that machinery when the task justifies it.

Routing is not the whole answer

Sending each task to the model best suited to it, by capability, governance and unit economics, covers most of the ground. Some tasks do not resolve into a single best model at all. A capable open-weight model can now do most of a piece of specialist work competently and cheaply, but not all of it, and the gap between “most” and “all” is exactly where the cost of a wrong answer lives. Rather than escalating the whole task to a frontier model out of caution, the more precise answer is to combine models: one doing the work, another checking it, or several attempting it so a separate process can judge between them.

This is a genuinely new option rather than a repackaged one. Z.AI’s GLM 5.2 was, by mid-2026, the first open-weight model to clear the capability level that had defined the start of the agentic era, which is what makes pairing it with a stronger adviser a serious architecture rather than a compromise dressed up as one.

A worker and an adviser

The simplest orchestration pattern pairs a cheaper model doing the bulk of the work with a stronger model available as an adviser. Legal technology has already shown the shape of this: reported systems from Harvey and Fireworks pair an open-weight GLM worker with an Opus-class adviser for legal drafting and review, running the cheaper model as the default and calling the stronger one in on specific, defined conditions rather than for every request.

The worker's remitWhat the cheaper model is trusted to finish alone
Defined by task type and a confidence or complexity signal, not left to the model's own judgement about when it is unsure.
The escalation triggerWhat calls the adviser in
A specific, testable condition: a flagged ambiguity, a case outside the worker's demonstrated competence, or a check the worker itself cannot perform on its own output.
The resolution ruleWhat happens when worker and adviser disagree
Decided in advance and made visible in the output, rather than silently defaulting to whichever model spoke last or is assumed to be stronger.

None of this removes the need for the underlying discipline set out in how we choose and govern models per task: the worker and the adviser are each still selected on capability, governance and unit economics. The pairing is an additional design decision on top of that choice, not a replacement for it.

A panel, a judge, a synthesiser

A second pattern goes further: several models attempt the same task independently, a judge, itself a model or a defined evaluation process, scores or filters the attempts against stated criteria, and a synthesiser produces the final answer from what the judge selected or combined. Systems built on this shape, such as OpenRouter’s Fusion, aim to reach higher capability than any single model in the panel, at a lower cost than always calling the most expensive one.

Building this well means treating the judge as seriously as the panel. A judge is itself a model with its own biases, and an unexamined one will tend to reward answers that resemble its own training rather than answers that are actually correct. So the judge’s criteria have to be written down and testable, its agreement with a trusted reference checked before it is trusted to decide anything, and its decisions logged so a pattern of questionable judging is visible rather than buried inside an average.

What can go wrong

The failure modes are specific enough to design against. Hand-off points between a worker and an adviser, or between a panel and its judge, are where context is most often lost: an adviser reviewing only the worker’s output, without the reasoning or the source material behind it, is reviewing a summary rather than the work. A judge that has not been checked against a trusted reference can reward confident phrasing over correct content, which is a worse failure than an obviously wrong answer because it is not visible without deliberate evaluation. And the additional cost of running several models is real: orchestration that is not saving more than it spends, in money, latency or error rate, is complexity for its own sake.

When one call is still right

Most tasks do not need any of this. Where a single model, correctly chosen, reliably does the job to the standard the work requires, adding a worker-adviser pairing or a panel-judge-synthesiser pipeline is added cost and added failure surface for no measured benefit. Orchestration earns its place on tasks with a real accuracy gap between the cheap answer and the trustworthy one, high enough stakes that closing that gap is worth the added machinery, and enough volume that the pattern is worth building and maintaining rather than escalating the occasional hard case by hand.

Bring us the task sitting between two models.

Describe the work that one model does well enough most of the time, and badly the rest of it. We will show you what a worker-adviser or panel-judge pattern would look like for it, and whether it is worth building.

Your message goes to the people who would do the work.