Model routing is three decisions, made on your own data: 1. Model pool (GPT-5 nano and GPT-5), 2. Classification (SIMPLE to nano, COMPLEX to GPT-5), 3. Scope (per conversation or per question).

Model Routing, Part 2: Choosing the Scope

Objective

In Part 1, I compared ways to choose which model should answer a support request. This post continues that experiment with a second decision: how long should that model choice last—a whole conversation or just one question? The small pilot illustrates a method, not a general recommendation.

I compared four policies on six new three-question conversations: always GPT-5 nano, always GPT-5, per-conversation routing and per-question routing. Nano and both routing policies got 4 of 6 conversations fully correct; GPT-5 got all six. Routing cost about nine times nano, without improving the fully-correct-conversation count. Neither scope clearly won.

The Agent and Classifier

The agent answers marketplace support questions using read-only SQL and policy search, through the OpenAI Agents SDK and LiteLLM, with Langfuse tracing. Part 1 compared classifiers on 18 cases. Here I reuse its domain-aware nano classifier: SIMPLE → nano, COMPLEX → GPT-5.

Three boxes: 1. Model pool (a fixed, approved set of models; ours are GPT-5 nano and GPT-5), 2. Classification (label each request, then map each label to a model; ours map SIMPLE to nano and COMPLEX to GPT-5), 3. Scope (per conversation, per question, per model call, named stage). Below, one box lists what a gateway such as LiteLLM can manage: allowed models, built-in classifiers, per-conversation selection through session affinity, per-question selection through user_turn and per-model-call selection through every_request. Another box lists what the application must supply: session identity, structured messages and named stages. A note says that in the conversation-granularity experiment the author's own runner did the classification, selection and pinning, and LiteLLM forwarded each request to the selected model.
Part 1 covered the model pool and classification. This post covers the third decision, scope.

My Starting Expectation

I expected per-conversation routing to preserve context and cache warmth, while per-question routing could save money on easy follow-ups. I also assumed switching would lose context or start cold. These were assumptions to test, not established properties.

Ways to Scope a Routing Decision

A gateway can manage most of these for you. LiteLLM’s documentation describes three built-in options:

  • Per conversation: session affinity pins the model chosen on the first turn of a session and skips reclassification on later turns.
  • Per question: classification_mode: user_turn classifies only new user questions, and carries that decision through the tool calls and continuation turns within the question.
  • Per model call: classification_mode: every_request (the default) classifies each request the agent sends to the gateway.

For any of these to work, the application supplies the session identity (LiteLLM reads a session_id from the request metadata) and well-structured messages, so the gateway can tell a new user question from a tool result.

Two more options need the application’s knowledge:

  • Named stages: a deliberate rule such as “nano gathers the evidence, GPT-5 writes the final answer”. The gateway can reclassify every call, but only the application knows which call is “gathering” and which is “writing”.
  • Escalation: start on nano and switch up when a check fails. This needs a reliable signal that an answer is wrong.

Here is what a single question looks like inside the agent:

Illustrative sequence for one question, "Why hasn't Cedar been paid for order-1?", in five steps: model call 1 decides to query payout records; a tool runs read-only SQL; model call 2 reads the rows and decides to search the payout policy; a tool searches the policy documents; model call 3 writes the answer with citations. Notes below explain that per-question routing uses one model for all three calls (LiteLLM user_turn), automatic per-call routing reclassifies each call (every_request), and deliberate stage selection is decided by the application.
An illustrative question. Automatic per-call routing reclassifies each call; deliberate stage selection needs the application to name the stages.

I tested per conversation and per question, against the two fixed models.

How I Tested It

What I implemented. My own runner did the classification, the model selection and the pinning. LiteLLM simply forwarded each request to the model the runner had explicitly selected. So these results test the two routing policies, not LiteLLM’s native session-affinity or user_turn implementations. Every internal model call for a question used the model chosen for that question.

The conversations. I wrote 6 new conversation scenarios of 3 questions each, with expected answers written in advance. They cover:

  • simple questions leading to hard ones, and the reverse
  • a clarification
  • a purchase that changes the data mid-conversation
  • a clock change that lets a payout hold pass
  • simple lookups only

The policies. Each scenario ran once under four policies:

  • always nano
  • always GPT-5
  • per-conversation routing
  • per-question routing, where the classifier also sees the earlier conversation

Six scenarios × three questions × four policies is 72 answers, but only 6 distinct scenarios, not 72 independent tests. I couldn’t reuse saved answers as in Part 1, because each answer shapes the next question’s context.

Grading. An LLM reviewer, GPT-6 Astra, reviewed every answer against the expected facts, and I adjudicated the borderline cases. Three borderline partial grades were flagged; I reviewed and accepted them. A conversation counts as fully correct only if all three of its answers are correct, because a support conversation with one bad answer is still a failed conversation.

One Conversation, Two Scopes

Here is one scenario, “simple lookup to payout investigation”, under the two routing policies:

#QuestionPer conversationPer question
1“For order-2, list the purchased product, quantity and seller.”nano (pinned for the conversation) ✅nano ✅
2“For that same order, explain why its seller has not received the payout and what must happen before release.”nano ⚠️ partialGPT-5 ✅
3“What exact amount is currently payable to that seller for this order, after the marketplace fee?”nano ✅ ₹1,618.20GPT-5 ✅ ₹1,618.20

The first question looked simple, so per-conversation routing pinned the whole conversation to nano, including the harder second question. Per-question routing classified question 2 as COMPLEX and sent it to GPT-5. The classifier’s reason: “Determining why a seller payout is not completed requires tracing allocations, payments, shipments, returns/refunds, and payout status, plus policy-based eligibility.”

The Results

The headline measure is conversations with all three answers correct. Costs are cache-adjusted usage estimates for these newly generated answers. They are reconstructed from the token and cached-token counts the provider reported, at fixed list rates.

PolicyConversations fully correctCache-adjusted usage estimate, 6 conversationsvs always GPT-5
Always nano4 / 6$0.014−94.9%
Always GPT-56 / 6$0.271—
Per conversation4 / 6$0.126−53.4%
Per question4 / 6$0.127−53.2%
Two bar charts titled "Same fully-correct-conversation count for nano and both routing policies; 6 conversation scenarios, one pass each". Left: conversations with all three answers correct: always nano 4 of 6, always GPT-5 6 of 6, per conversation 4 of 6, per question 4 of 6. Right: cache-adjusted usage estimate for six conversations: always nano $0.014, always GPT-5 $0.271, per conversation $0.126, per question $0.127.
Nano and both routing policies had the same fully-correct-conversation count; only GPT-5 alone got all six. Six scenarios, one pass each.
  • Nano and both routing policies each got 4 of 6 conversations fully correct; GPT-5 got 6 of 6.
  • Routing cost about nine times nano, and it did not improve the fully-correct-conversation count.
  • Routing did reduce incorrect individual answers. Out of 18 answers per policy, nano got 15 correct, 1 partial and 2 incorrect; both routing policies got 15 / 2 / 1. Fewer wrong answers did not produce more fully correct conversations.
  • Neither granularity established a compelling advantage. Per conversation and per question got the same conversation score at almost the same cost.

Both routing policies chose nano for 12 of the 18 questions and GPT-5 for 6, though not always the same questions. Per-question routing switched models four times across the six conversations.

Right vs Wrong: Two Conversations in Detail

1. Per-question routing caught the harder question. In the scenario above, question 2 asks why the seller hasn’t been paid. The correct answer is that payment is captured, the shipment is still pending, and the payout is held until delivery is confirmed and a 24-hour hold has passed.

  • Per-conversation routing (pinned to nano): said the payout “is currently held” until “delivery/fulfillment conditions are satisfied”. It left out the 24-hour hold, so it was graded partial.
  • Per-question routing (sent to GPT-5): “Policy requires delivery confirmation plus a 24-hour hold… No delivery confirmation is recorded yet.” Correct.

Routing the hard question to the stronger model helped here. But there is a twist: in the separate always-nano run, nano answered this same question correctly. Same model, same question, different run, different grade. One run is not proof.

2. When every question looks simple, no scope helps. This scenario was about a name that means two things: “Cedar” is both a seller (Cedar & Co.) and a product (Cedar Desk).

  • “Please list purchases associated with Cedar.” Nano correctly asked whether I meant the seller or the product.
  • “The seller Cedar & Co., not the desk product.” Nano invented a “today only” date filter nobody asked for and replied that there were no purchases. In fact there was one: order-1, with two Canvas Weekenders. Graded incorrect.
  • “Within those purchases, show only items that have not been delivered.” Nano answered “none”. That happens to be the right result, since the order had been delivered, but only because the previous answer had filtered everything out. Right answer, wrong reason, graded partial.

The classifier labeled all three questions SIMPLE, and on the surface they are: a list, a clarification, a filter. So both routing policies used nano throughout and made the same mistakes. GPT-5 alone got all three right. Per-question routing only helps if the classifier can tell that a question is risky, and follow-ups often look simpler than they are. Once a wrong answer enters the conversation, later answers build on it.

What It Costs

The pilot also recorded latency and caching. Full-price estimates are shown next to the cache-adjusted estimates so the effect of caching is visible.

PolicyCache-adjusted estimateFull-price estimateCached share of inputLatency p50 / p95Model switches
Always nano$0.014$0.02050.6%10.3 s / 20.2 s0
Always GPT-5$0.271$0.46064.0%13.2 s / 34.8 s0
Per conversation$0.126$0.20654.0%10.4 s / 33.1 s0
Per question$0.127$0.21957.9%13.4 s / 38.8 s4
  • Caching changes the estimate a lot. For always-GPT-5, the cache-adjusted usage estimate was about 40% below the full-price estimate. Neither figure is an invoice.
  • Report cache differences, don’t explain them away. Per-question routing switched models four times, and its cached share was 57.9% against 54.0% for per-conversation. I am not attributing that difference to model switching, or to anything else, because two things make it hard to interpret:
    • Provider caches were shared and could already be warm across runs.
    • One conversation was interrupted mid-run when my API credit ran out. It resumed after funding, which may have changed cache warmth.
  • Treat the latency figures as rough. p95 over 18 answers per policy is an indicator, not a production estimate.

What Happened to My Expectation

I expected per-conversation routing to win on context and caching, with per-question routing trading some of both for lower cost. Here is how each part held up:

  • Context: my assumption was wrong. A model switch doesn’t lose the conversation. The application passes the conversation history with every call, so the model answering question 2 sees question 1 and its answer, whichever model it is. What carries over from one question to the next is the content of earlier answers, including their mistakes, not the model.
  • Caching: I didn’t observe the penalty I expected. The per-question policy’s cached share was higher, not lower, and that held within each model too: nano 51.5% versus 47.8%, GPT-5 66.9% versus 61.6%. With shared cache warmth, the funding interruption and a single pass, this observation doesn’t tell me why.
  • Cost: effectively the same. $0.126 for per-conversation, $0.127 for per-question.
  • Accuracy: the same conversation score, with different mistakes. Per-question routing rescued one question the pinned conversation got partly wrong. In another conversation, neither scope helped, because every question looked simple.

So the pilot didn’t confirm my expectation, and it didn’t reverse it either. It was inconclusive. The useful part is seeing that the differences between the scopes were smaller and more situational than I assumed.

Pros and Cons of Each Scope

No scope is best in general. Each one trades something:

ScopeProsConsConsider it when
Per conversationOne classification per conversation; one model; simplest to run and explainThe first question decides for all later ones; a simple opener can leave a hard follow-up on nanoConversations tend to stay at one level of difficulty
Per questionAdapts when a conversation gets harder or easier; conversation history still carries overA classification on every question; follow-ups often look simpler than they are; more model switches to monitorConversations often mix quick lookups with investigations
Per model call (automatic)The gateway can reclassify every call with no application changesCalls within one question may land on different models; more classifications; the classifier sees the call, not the stageYou want fine-grained routing without writing stage logic
Named stages (application-defined)Can reserve the stronger model for the stage that needs it, such as writing the final answerThe application must identify the stages; the final stage often carries the most context, so it may not save muchFailures cluster in one stage, such as interpreting evidence
EscalationPays for the stronger model only when the cheaper one failsNeeds a reliable signal that an answer is wrong; fluent wrong answers slip through; slower when it escalatesYou have strong automatic checks on answers

A Checklist for Choosing Routing

Steps 1–4 are from Part 1; steps 5 and 6 come from this post.

  1. Shortlist models on benchmarks, domain fit, cost and latency, from approved providers that support your agent’s needs.
  2. Run evals on your golden set, with single questions and full conversations, to narrow the shortlist. Then fix the pool and pin the versions.
  3. Weigh cost savings against accuracy. If the cheaper model meets your accuracy bar, use it. If only the stronger one does and its cost is acceptable, use it. Route only when routing saves meaningful cost while keeping accuracy within your bar.
  4. Design the classification. Choose labels that matter for your traffic and map them to models. Start with keyword rules, then try semantic or LLM-based classifiers built from real user journeys.
  5. Choose the granularity on real conversations. Compare per conversation and per question (and per call or named stages if they fit your agent) on multi-question conversations, watching conversation-level correctness, cost, latency and caching.
  6. Validate and monitor. Re-test on fresh scenarios and repeat unstable cases before deciding, and rerun your evals whenever a model changes or retires.

Limits and What’s Next

This was a small pilot, and I would not treat it as a benchmark:

  • Sample size: 6 conversation scenarios of 3 questions, one pass per policy.
  • Grading: an LLM reviewer plus my adjudication of borderline cases.
  • Classifier: a prompt developed on known cases.
  • Implementation: my own runner, not the gateway’s native routing.
  • Caching: shared cache warmth and a funding interruption; the figures are observations, not causal proof.

The “different run, different grade” twist above shows how much one pass can vary. Next, I want to:

  • run more conversations, with repeated passes
  • repeat the test with LiteLLM’s native session affinity and user_turn
  • try named-stage routing, with nano gathering evidence and GPT-5 writing the answer
  • test escalation, once I have a reliable signal that an answer is wrong
  • run controlled cache tests

Acknowledgements

This post was written with the help of Claude Opus 5.5 in Claude Code, which also helped me research the routing landscape and plan the experiments. The application and experiments were built and run with GPT-6 Astra, which also reviewed answers against the expected results.

References

This Series

Earlier Posts

Routing Tools

  • LiteLLM Auto Routing: session affinity, classification_mode (every_request and user_turn) and classifier configuration.

1 thought on “Model Routing, Part 2: Choosing the Scope”

Leave a comment