By Roland CadavosModel Routing
Model Routing: When Choosing the Model Became a Runtime Decision (2026)
By October 2026, picking a model was no longer a one-time config choice. Price cuts, speed tiers, deprecations, and multi-model orchestration turned it into a decision your system makes on every request.
By October 2026, “which model should we use?” had stopped being an architecture decision and become a runtime one. For two years, most teams picked a model, wrote its name into config, and revisited the choice only when something broke. The last weeks of September made that habit look expensive. Anthropic and OpenAI both shipped cheaper models, paid speed tiers multiplied, deprecation notices kept arriving, and GitHub began coordinating several models inside a single coding turn. The model was no longer a fixed dependency; it was a decision made per request.
The economics moved fastest. Anthropic released a model it said cost markedly less to run than the one it replaced, and sold a faster mode of it at a premium. At DevDay, OpenAI introduced a mid-tier model it described as near-flagship intelligence for a fifth of the price, alongside a paid speed tier for work where latency matters most. OpenAI’s own guide to its new model family put the new default plainly: balance capability, cost, and latency by choosing the model, reasoning effort, and speed that fit each task.
GitHub turned the idea into a product. Its HydraFusion research preview, which reached VS Code and the Copilot app at the end of September, does not pick one model for a task; it picks a workflow. A simple task goes to a single model. In a cascade, an efficient model drafts and a quality gate either accepts the result or escalates to a stronger one. In a critique, a read-only reviewer from a different model family reviews the draft and the author revises once. GitHub reported that, in offline evaluations, the approach matched or came close to a frontier baseline at a much lower estimated cost.
The primitives followed. OpenAI’s Decisions API, in limited preview, uses its most efficient model to classify inputs, route requests, or choose an agent’s next action from a fixed set of answers. That is routing sold as a feature: a cheap, fast model deciding where the expensive work should go. For application developers, the architectural consequence is plain. The model identifier no longer belongs scattered through business logic; it belongs behind one interface that takes the task, a quality bar, and a budget, and decides how to spend them.
Routing only works if you can tell when cheap was good enough, which is where June’s argument about evals came back with interest. A cascade is only as trustworthy as its quality gate, and a quality gate is an eval running in production: tests that pass, schemas that validate, a reviewer that agrees. GitHub built its orchestration around exactly that, accounting for token cost across every draft, critique, retry, and escalation, and rejecting patches that fail validation. Without measurable acceptance criteria, a team cannot route safely; it can only guess which model to trust.
The churn cut the other way, too. On the last day of September, Anthropic notified developers that one of its older models would be retired from its API by the end of November, pointing them to a replacement; two days later, GitHub deprecated a batch of models in Copilot. Every hard-coded model string is a dependency with an expiry date. The sane response is to treat model versions like any other pinned dependency: declared in one place, covered by a regression suite, and upgraded on purpose rather than in a hurry.
Routing has costs of its own. Every extra hop adds latency, and a cascade that escalates too often can cost more than calling the strong model directly. Mixing providers multiplies the surfaces to secure, monitor, and contract with, and a critic from another model family only helps if someone acts on what it says. Debugging gets harder as well: when an answer came from a draft, a critique, and a revision, a bad result has three suspects. The fix is the observability you would want for any distributed system—trace every leg, log every routing decision, and keep the budget visible.
The niche takeaway for working developers: stop treating the model as a constant and start treating it as a decision your system makes. Put model choice behind one interface, define what good enough means for each task, measure it, and let cheaper models handle what they can while a gate decides when to escalate. Pin versions, watch the deprecation pages, and keep a regression suite ready for the next swap. In 2026, the teams that routed on evidence got close to frontier results without frontier bills; everyone else paid flagship prices for work a smaller model could have done.