All posts
model-routinglitellmself-hosted-aicost-controloperations

Automatic Model Routing: What to Test Before Using It

I assess automatic model routing by task quality, full cost, fallback behaviour and data permissions. A complexity estimate must not decide where data may go.

Stéphane Lepain··Updated ·4 min read

Updated 13 September 2026: this article now presents a general method. Examples are illustrative, not customer results.

An automatic model choice can simplify an assistant's interface. I assess whether it helps the task before treating it as the default: the result must meet the same quality and data-access requirements, at an acceptable full cost.

Model routing must respect permitted data destinations before considering task quality, usage cost or fallback behaviour. A harder question does not authorize an external route. If no permitted route meets the task's requirements, the workflow needs to stop or hand the task to the named reviewer.

This article describes a design and test method. It does not publish a private configuration, measured saving or customer deployment.

Why model dropdowns are an operations problem

A person trying to finish a task may not know which model to choose. A router can make that choice according to an explicit policy, but it also adds decisions that an operator needs to understand and test.

The policy might use a fixed task mapping, a heuristic or another classification step. Each has costs and failure modes. Verify the capabilities of the installed software rather than assuming all routers behave alike.

What the router actually does

Write down the input to the decision, the permitted destinations, fallback behaviour and what happens when the first attempt fails. Distinguish routing an entire conversation from choosing again on each message.

The public LiteLLM routing documentation describes available mechanisms. A documentation example is not evidence that a particular deployment is configured correctly.

"Automatic" does not mean optimal

Use fictional test conversations that change difficulty: a short greeting followed by a complex task, and a complex task followed by a simple reply. Observe whether the model choice changes and whether that helps or harms the result.

Keeping a conversation on one model may improve continuity or prefix reuse, but it can also retain an unsuitable choice. Switching models can introduce other costs. Measure the actual result rather than declaring either approach universally better. If caching is part of the cost argument, verify it against the provider's reported usage. A setting in a file is not evidence that the request used it.

Complexity is not a data-classification policy

Difficulty and permission are separate decisions. A harder question must not authorise sending data to a new provider.

I first agree which information may reach which services. Any automatic choice must stay inside that boundary, including retries and fallbacks. Test prohibited destinations and failure cases as well as the expected route.

What I will measure next

For a proposed routing policy, I measure:

  • whether the completed result meets the agreed checks;
  • which routes were actually used, including fallbacks;
  • API cost, classification overhead and failed attempts;
  • remaining human review and correction;
  • behaviour when the task changes during a conversation.

There is no claimed saving until the comparison supports it. A cheaper selected model can still produce a more expensive completed task.

Questions people ask about automatic model routing

Does routing add latency?

It depends on the implementation. Measure the decision step and the full response, including retries.

Why keep a thread on one model?

Continuity and prefix reuse may be useful, but a harder follow-up may need a different approach. Test representative conversations.

Can routing change the data boundary?

It must not change the agreed permissions. Verify every possible destination and fallback in the actual configuration.

Is automatic routing required for an assistant?

No. A fixed model or explicit task mapping may be enough. I choose the simpler approach when it meets the task's requirements.

I scope useful assistance and automation around the work, human checks and full cost. The evidence page explains the public examples and their limits.