Tag Archives: LiteLLM

Model routing is three decisions, made on your own data: 1. Model pool (GPT-5 nano and GPT-5), 2. Classification (SIMPLE to nano, COMPLEX to GPT-5), 3. Scope (per conversation or per question).

Model Routing, Part 2: Choosing the Scope

Objective

In Part 1, I compared ways to choose which model should answer a support request. This post continues that experiment with a second decision: how long should that model choice last—a whole conversation or just one question? The small pilot illustrates a method, not a general recommendation.

I compared four policies on six new three-question conversations: always GPT-5 nano, always GPT-5, per-conversation routing and per-question routing. Nano and both routing policies got 4 of 6 conversations fully correct; GPT-5 got all six. Routing cost about nine times nano, without improving the fully-correct-conversation count. Neither scope clearly won.

The Agent and Classifier

The agent answers marketplace support questions using read-only SQL and policy search, through the OpenAI Agents SDK and LiteLLM, with Langfuse tracing. Part 1 compared classifiers on 18 cases. Here I reuse its domain-aware nano classifier: SIMPLE → nano, COMPLEX → GPT-5.

Three boxes: 1. Model pool (a fixed, approved set of models; ours are GPT-5 nano and GPT-5), 2. Classification (label each request, then map each label to a model; ours map SIMPLE to nano and COMPLEX to GPT-5), 3. Scope (per conversation, per question, per model call, named stage). Below, one box lists what a gateway such as LiteLLM can manage: allowed models, built-in classifiers, per-conversation selection through session affinity, per-question selection through user_turn and per-model-call selection through every_request. Another box lists what the application must supply: session identity, structured messages and named stages. A note says that in the conversation-granularity experiment the author's own runner did the classification, selection and pinning, and LiteLLM forwarded each request to the selected model.
Part 1 covered the model pool and classification. This post covers the third decision, scope.

My Starting Expectation

I expected per-conversation routing to preserve context and cache warmth, while per-question routing could save money on easy follow-ups. I also assumed switching would lose context or start cold. These were assumptions to test, not established properties.

Ways to Scope a Routing Decision

A gateway can manage most of these for you. LiteLLM’s documentation describes three built-in options:

  • Per conversation: session affinity pins the model chosen on the first turn of a session and skips reclassification on later turns.
  • Per question: classification_mode: user_turn classifies only new user questions, and carries that decision through the tool calls and continuation turns within the question.
  • Per model call: classification_mode: every_request (the default) classifies each request the agent sends to the gateway.

For any of these to work, the application supplies the session identity (LiteLLM reads a session_id from the request metadata) and well-structured messages, so the gateway can tell a new user question from a tool result.

Two more options need the application’s knowledge:

  • Named stages: a deliberate rule such as “nano gathers the evidence, GPT-5 writes the final answer”. The gateway can reclassify every call, but only the application knows which call is “gathering” and which is “writing”.
  • Escalation: start on nano and switch up when a check fails. This needs a reliable signal that an answer is wrong.

Here is what a single question looks like inside the agent:

Illustrative sequence for one question, "Why hasn't Cedar been paid for order-1?", in five steps: model call 1 decides to query payout records; a tool runs read-only SQL; model call 2 reads the rows and decides to search the payout policy; a tool searches the policy documents; model call 3 writes the answer with citations. Notes below explain that per-question routing uses one model for all three calls (LiteLLM user_turn), automatic per-call routing reclassifies each call (every_request), and deliberate stage selection is decided by the application.
An illustrative question. Automatic per-call routing reclassifies each call; deliberate stage selection needs the application to name the stages.

I tested per conversation and per question, against the two fixed models.

How I Tested It

What I implemented. My own runner did the classification, the model selection and the pinning. LiteLLM simply forwarded each request to the model the runner had explicitly selected. So these results test the two routing policies, not LiteLLM’s native session-affinity or user_turn implementations. Every internal model call for a question used the model chosen for that question.

The conversations. I wrote 6 new conversation scenarios of 3 questions each, with expected answers written in advance. They cover:

  • simple questions leading to hard ones, and the reverse
  • a clarification
  • a purchase that changes the data mid-conversation
  • a clock change that lets a payout hold pass
  • simple lookups only

The policies. Each scenario ran once under four policies:

  • always nano
  • always GPT-5
  • per-conversation routing
  • per-question routing, where the classifier also sees the earlier conversation

Six scenarios × three questions × four policies is 72 answers, but only 6 distinct scenarios, not 72 independent tests. I couldn’t reuse saved answers as in Part 1, because each answer shapes the next question’s context.

Grading. An LLM reviewer, GPT-6 Astra, reviewed every answer against the expected facts, and I adjudicated the borderline cases. Three borderline partial grades were flagged; I reviewed and accepted them. A conversation counts as fully correct only if all three of its answers are correct, because a support conversation with one bad answer is still a failed conversation.

One Conversation, Two Scopes

Here is one scenario, “simple lookup to payout investigation”, under the two routing policies:

#QuestionPer conversationPer question
1“For order-2, list the purchased product, quantity and seller.”nano (pinned for the conversation) ✅nano ✅
2“For that same order, explain why its seller has not received the payout and what must happen before release.”nano ⚠️ partialGPT-5 ✅
3“What exact amount is currently payable to that seller for this order, after the marketplace fee?”nano ✅ ₹1,618.20GPT-5 ✅ ₹1,618.20

The first question looked simple, so per-conversation routing pinned the whole conversation to nano, including the harder second question. Per-question routing classified question 2 as COMPLEX and sent it to GPT-5. The classifier’s reason: “Determining why a seller payout is not completed requires tracing allocations, payments, shipments, returns/refunds, and payout status, plus policy-based eligibility.”

The Results

The headline measure is conversations with all three answers correct. Costs are cache-adjusted usage estimates for these newly generated answers. They are reconstructed from the token and cached-token counts the provider reported, at fixed list rates.

PolicyConversations fully correctCache-adjusted usage estimate, 6 conversationsvs always GPT-5
Always nano4 / 6$0.014−94.9%
Always GPT-56 / 6$0.271—
Per conversation4 / 6$0.126−53.4%
Per question4 / 6$0.127−53.2%
Two bar charts titled "Same fully-correct-conversation count for nano and both routing policies; 6 conversation scenarios, one pass each". Left: conversations with all three answers correct: always nano 4 of 6, always GPT-5 6 of 6, per conversation 4 of 6, per question 4 of 6. Right: cache-adjusted usage estimate for six conversations: always nano $0.014, always GPT-5 $0.271, per conversation $0.126, per question $0.127.
Nano and both routing policies had the same fully-correct-conversation count; only GPT-5 alone got all six. Six scenarios, one pass each.
  • Nano and both routing policies each got 4 of 6 conversations fully correct; GPT-5 got 6 of 6.
  • Routing cost about nine times nano, and it did not improve the fully-correct-conversation count.
  • Routing did reduce incorrect individual answers. Out of 18 answers per policy, nano got 15 correct, 1 partial and 2 incorrect; both routing policies got 15 / 2 / 1. Fewer wrong answers did not produce more fully correct conversations.
  • Neither granularity established a compelling advantage. Per conversation and per question got the same conversation score at almost the same cost.

Both routing policies chose nano for 12 of the 18 questions and GPT-5 for 6, though not always the same questions. Per-question routing switched models four times across the six conversations.

Right vs Wrong: Two Conversations in Detail

1. Per-question routing caught the harder question. In the scenario above, question 2 asks why the seller hasn’t been paid. The correct answer is that payment is captured, the shipment is still pending, and the payout is held until delivery is confirmed and a 24-hour hold has passed.

  • Per-conversation routing (pinned to nano): said the payout “is currently held” until “delivery/fulfillment conditions are satisfied”. It left out the 24-hour hold, so it was graded partial.
  • Per-question routing (sent to GPT-5): “Policy requires delivery confirmation plus a 24-hour hold… No delivery confirmation is recorded yet.” Correct.

Routing the hard question to the stronger model helped here. But there is a twist: in the separate always-nano run, nano answered this same question correctly. Same model, same question, different run, different grade. One run is not proof.

2. When every question looks simple, no scope helps. This scenario was about a name that means two things: “Cedar” is both a seller (Cedar & Co.) and a product (Cedar Desk).

  • “Please list purchases associated with Cedar.” Nano correctly asked whether I meant the seller or the product.
  • “The seller Cedar & Co., not the desk product.” Nano invented a “today only” date filter nobody asked for and replied that there were no purchases. In fact there was one: order-1, with two Canvas Weekenders. Graded incorrect.
  • “Within those purchases, show only items that have not been delivered.” Nano answered “none”. That happens to be the right result, since the order had been delivered, but only because the previous answer had filtered everything out. Right answer, wrong reason, graded partial.

The classifier labeled all three questions SIMPLE, and on the surface they are: a list, a clarification, a filter. So both routing policies used nano throughout and made the same mistakes. GPT-5 alone got all three right. Per-question routing only helps if the classifier can tell that a question is risky, and follow-ups often look simpler than they are. Once a wrong answer enters the conversation, later answers build on it.

What It Costs

The pilot also recorded latency and caching. Full-price estimates are shown next to the cache-adjusted estimates so the effect of caching is visible.

PolicyCache-adjusted estimateFull-price estimateCached share of inputLatency p50 / p95Model switches
Always nano$0.014$0.02050.6%10.3 s / 20.2 s0
Always GPT-5$0.271$0.46064.0%13.2 s / 34.8 s0
Per conversation$0.126$0.20654.0%10.4 s / 33.1 s0
Per question$0.127$0.21957.9%13.4 s / 38.8 s4
  • Caching changes the estimate a lot. For always-GPT-5, the cache-adjusted usage estimate was about 40% below the full-price estimate. Neither figure is an invoice.
  • Report cache differences, don’t explain them away. Per-question routing switched models four times, and its cached share was 57.9% against 54.0% for per-conversation. I am not attributing that difference to model switching, or to anything else, because two things make it hard to interpret:
    • Provider caches were shared and could already be warm across runs.
    • One conversation was interrupted mid-run when my API credit ran out. It resumed after funding, which may have changed cache warmth.
  • Treat the latency figures as rough. p95 over 18 answers per policy is an indicator, not a production estimate.

What Happened to My Expectation

I expected per-conversation routing to win on context and caching, with per-question routing trading some of both for lower cost. Here is how each part held up:

  • Context: my assumption was wrong. A model switch doesn’t lose the conversation. The application passes the conversation history with every call, so the model answering question 2 sees question 1 and its answer, whichever model it is. What carries over from one question to the next is the content of earlier answers, including their mistakes, not the model.
  • Caching: I didn’t observe the penalty I expected. The per-question policy’s cached share was higher, not lower, and that held within each model too: nano 51.5% versus 47.8%, GPT-5 66.9% versus 61.6%. With shared cache warmth, the funding interruption and a single pass, this observation doesn’t tell me why.
  • Cost: effectively the same. $0.126 for per-conversation, $0.127 for per-question.
  • Accuracy: the same conversation score, with different mistakes. Per-question routing rescued one question the pinned conversation got partly wrong. In another conversation, neither scope helped, because every question looked simple.

So the pilot didn’t confirm my expectation, and it didn’t reverse it either. It was inconclusive. The useful part is seeing that the differences between the scopes were smaller and more situational than I assumed.

Pros and Cons of Each Scope

No scope is best in general. Each one trades something:

ScopeProsConsConsider it when
Per conversationOne classification per conversation; one model; simplest to run and explainThe first question decides for all later ones; a simple opener can leave a hard follow-up on nanoConversations tend to stay at one level of difficulty
Per questionAdapts when a conversation gets harder or easier; conversation history still carries overA classification on every question; follow-ups often look simpler than they are; more model switches to monitorConversations often mix quick lookups with investigations
Per model call (automatic)The gateway can reclassify every call with no application changesCalls within one question may land on different models; more classifications; the classifier sees the call, not the stageYou want fine-grained routing without writing stage logic
Named stages (application-defined)Can reserve the stronger model for the stage that needs it, such as writing the final answerThe application must identify the stages; the final stage often carries the most context, so it may not save muchFailures cluster in one stage, such as interpreting evidence
EscalationPays for the stronger model only when the cheaper one failsNeeds a reliable signal that an answer is wrong; fluent wrong answers slip through; slower when it escalatesYou have strong automatic checks on answers

A Checklist for Choosing Routing

Steps 1–4 are from Part 1; steps 5 and 6 come from this post.

  1. Shortlist models on benchmarks, domain fit, cost and latency, from approved providers that support your agent’s needs.
  2. Run evals on your golden set, with single questions and full conversations, to narrow the shortlist. Then fix the pool and pin the versions.
  3. Weigh cost savings against accuracy. If the cheaper model meets your accuracy bar, use it. If only the stronger one does and its cost is acceptable, use it. Route only when routing saves meaningful cost while keeping accuracy within your bar.
  4. Design the classification. Choose labels that matter for your traffic and map them to models. Start with keyword rules, then try semantic or LLM-based classifiers built from real user journeys.
  5. Choose the granularity on real conversations. Compare per conversation and per question (and per call or named stages if they fit your agent) on multi-question conversations, watching conversation-level correctness, cost, latency and caching.
  6. Validate and monitor. Re-test on fresh scenarios and repeat unstable cases before deciding, and rerun your evals whenever a model changes or retires.

Limits and What’s Next

This was a small pilot, and I would not treat it as a benchmark:

  • Sample size: 6 conversation scenarios of 3 questions, one pass per policy.
  • Grading: an LLM reviewer plus my adjudication of borderline cases.
  • Classifier: a prompt developed on known cases.
  • Implementation: my own runner, not the gateway’s native routing.
  • Caching: shared cache warmth and a funding interruption; the figures are observations, not causal proof.

The “different run, different grade” twist above shows how much one pass can vary. Next, I want to:

  • run more conversations, with repeated passes
  • repeat the test with LiteLLM’s native session affinity and user_turn
  • try named-stage routing, with nano gathering evidence and GPT-5 writing the answer
  • test escalation, once I have a reliable signal that an answer is wrong
  • run controlled cache tests

Acknowledgements

This post was written with the help of Claude Opus 5.5 in Claude Code, which also helped me research the routing landscape and plan the experiments. The application and experiments were built and run with GPT-6 Astra, which also reviewed answers against the expected results.

References

This Series

Earlier Posts

Routing Tools

  • LiteLLM Auto Routing: session affinity, classification_mode (every_request and user_turn) and classifier configuration.
Model routing is three decisions, made on your own data: 1. Model pool (GPT-5 nano and GPT-5), 2. Classification (SIMPLE to nano, COMPLEX to GPT-5), 3. Scope (per conversation or per question).

Model Routing, Part 1: Choosing a Classifier

Objective

How do you choose a model router for your own application? I tested six classifiers on a marketplace support agent to compare cost against answer quality. This small study demonstrates the decision process, not a general ranking of products.

Both GPT-5 nano and GPT-5 answered 18 support cases: 16 single-turn and two two-turn cases, or 20 answer turns per model. The LLM classifier and JEV each selected fully correct answers for 88.9% of cases, versus 94.4% for GPT-5 alone and 50% for nano, at an estimated 55–58% of GPT-5’s cost.

Part 2 examines how long a routing decision should last across a conversation.

The Use Case

I built a simulated marketplace with buyers, sellers, orders, shipments, payouts, returns, refunds and versioned policies. Its internal support agent answers questions using read-only SQL and policy search, citing its evidence. It runs on the OpenAI Agents SDK, through LiteLLM, with Langfuse tracing; payments and business data are simulated.

The workload mixes quick lookups—“What products are in order-2?”—with investigations such as “Why hasn’t this seller been paid?” The latter can require payment records, delivery events, policy and a time calculation.

Routing Is Three Decisions

Model routing is not one setting. It is three decisions:

Three boxes: 1. Model pool (a fixed, approved set of models; ours are GPT-5 nano and GPT-5), 2. Classification (label each request, then map each label to a model; ours map SIMPLE to nano and COMPLEX to GPT-5), 3. Scope (per conversation, per question, per model call, named stage). Below, one box lists what a gateway such as LiteLLM can manage: allowed models, built-in classifiers, per-conversation selection through session affinity, per-question selection through user_turn and per-model-call selection through every_request. Another box lists what the application must supply: session identity, structured messages and named stages. A note says that in the conversation-granularity experiment the author's own runner did the classification, selection and pinning, and LiteLLM forwarded each request to the selected model.
Routing is three decisions. A gateway can manage most of them; the application supplies session identity and any named stages.
  1. Model pool: which approved models the router may choose between.
  2. Classification: how each request is labeled, and which model each label maps to.
  3. Scope: how long a routing decision lasts, such as a whole conversation, a single question or a single model call. That is Part 2.

Choosing and Fixing the Model Pool

Shortlist approved models on benchmarks, domain fit, cost and latency, then run evals on your own golden set. Candidates must support the agent’s tool calls and output formats. Cost per correct answer matters more than token price alone.

My evaluated pair was GPT-5 nano, the lower-cost candidate, and GPT-5, the stronger performer. Both use the same provider behind one gateway. I did not run public benchmarks myself; the choice rested on workload evals, cost and latency.

Keep routing within the evaluated, approved pool and pin model versions for reproducibility. Even OpenRouter Auto was restricted to these two candidates.

How I Checked the Quality

A router can only be judged as well as the grades behind it, so grading came first.

I built a golden set: cases with known business data behind them and expected answers written in advance. A case counts as fully correct only if every requested fact is right, in every turn of the case.

Grading combined human review and LLM review:

  • Calibration, earlier: in an earlier model comparison, I graded 36 answers by hand and had Claude Opus 5.5 grade them as an LLM judge. After clarifying the grading rules, the judge agreed with my grade on 31 / 36 (86%). That calibration applies to that judge on those answers.
  • This comparison: the nano grades are existing accepted grades. The GPT-5 answers were reviewed by an LLM reviewer, GPT-6 Astra, against the expected answers, and I adjudicated the borderline cases. The Claude calibration does not validate that reviewer; the adjudication is what anchors those grades.

Baselines: What Routing Has to Beat

First, the two models on their own over the 18 cases. The cases cover order lookups, policy questions (current and historical), payout eligibility, refund histories, incidents, reconciliation, an ambiguous name, and a request the agent must refuse.

ModelCases correct / partial / incorrectFully correctEstimated cost, all 18 cases
GPT-5 nano9 / 7 / 250.0%$0.023
GPT-517 / 1 / 094.4%$0.456

GPT-5 was far more accurate and cost about 20× more. That gap is what makes routing worth considering. The question routing has to answer is simple: how much cost can we save, and how much accuracy do we give up?

Classification: Labels Are a Design Choice

A classifier assigns a label to each request. A separate mapping turns that label into a model. Both are your choices.

  • Labels don’t have to be about difficulty. They can be task type (lookup, investigation, action request), business domain (payments, returns, fulfillment), risk, language, or a mix.
  • “Simple” doesn’t have to mean “cheapest”. A short but risky request such as “refund this order now” could be labeled simple and still be mapped to the strong model.
  • A pool can have more than two tiers.

For this study, I used the simplest useful setup for a two-model pool: SIMPLE → GPT-5 nano and COMPLEX → GPT-5, for the classifiers I configured. LiteLLM’s built-in tiers also include MEDIUM and REASONING; I mapped everything above SIMPLE to GPT-5. OpenRouter’s Auto Router is the exception: it does not use my labels at all (see below).

The Six Classifiers and How I Configured Them

Every rule, example and prompt below was written from the support-user journeys in my specification: the kinds of questions support staff actually ask. They were frozen before testing. The one exception is a separately versioned refinement of the semantic examples, which I made after seeing the first results. It is reported on its own and is not part of the results table.

How each one ran. For the first five, my own runner produced or collected the classification and chose the model. OpenRouter Auto made its own choice. None of this used a gateway’s automatic routing on live traffic.

  • Keyword rules, heuristic scorer and semantic similarity: LiteLLM’s native classification code, exercised separately. It was not routing live traffic through my gateway.
  • LLM classifier: GPT-5 nano called through my LiteLLM gateway.
  • JEV: called directly through OpenRouter’s Decisions API.
  • OpenRouter Auto: called through OpenRouter.

Keyword Rules

LiteLLM’s keyword rules map words to tiers without calling any model. My lists:

keyword_tier_rules:
- tier: SIMPLE
keywords: [list, show, quantity, stock, balance, price]
- tier: COMPLEX
keywords: [why, investigate, reconcile, history, policy, eligible, eligibility,
refund, failed, incident, cause, "can you", "can the", "can we"]

Simple words signal a direct lookup. Complex words signal explanation, investigation, policy, eligibility, money movement, or a request to act. If both lists match, the higher tier wins. If nothing matches, the question falls back to the heuristic scorer below.

Heuristic Scorer

This is LiteLLM’s default classifier, and I used its defaults unchanged on purpose, to see how a generic heuristic behaves out of the box. It scores seven text signals that were built with software questions in mind:

  • code words (weight 0.30)
  • reasoning phrases like “step by step” (0.25)
  • technical terms (0.25)
  • length (0.10)
  • simple phrases like “what is” (0.05)
  • “first… then” patterns (0.03)
  • several question marks (0.02)

LiteLLM lets you add your own technical keywords, which is how you would adapt it to a domain.

Semantic Similarity

LiteLLM can also match questions by meaning. I wrote 14 example sentences, embedded with text-embedding-3-small. A question takes the label of its closest example if the similarity is at least 0.5; otherwise it falls back to the heuristic. A few of the examples:

  • SIMPLE (6 examples): “Show the products purchased in this order and their quantities.” · “What is the listed price of this product?” · “List the sellers and their current outstanding amounts.”
  • COMPLEX (8 examples): “Explain why the buyer paid but the merchant has not received the money.” · “Determine whether delivery and the waiting period permit a payout now.” · “Can you approve a refund and transfer the money on my behalf?”

LLM Classifier

Here a small model, GPT-5 nano, reads a routing prompt and returns SIMPLE or COMPLEX with a one-line reason. The prompt has five parts:

  1. Purpose: classify the evidence and reasoning a correct answer needs, “not sentence length or how easy the answer sounds”.
  2. The application and its tools: read-only SQL, policy search, event history.
  3. “What makes business questions deceptively difficult”: eight domain traps. Examples: “Buyer payment, seller payout and buyer refund are different money movements”; “Current payout/refund rows can hide failed attempts preceding a success”; “Recorded state, eligibility and successful execution differ.”
  4. Decision rules: SIMPLE only for direct, bounded lookups. COMPLEX for causal explanation, event reconstruction, policy applied to live data, time-boundary calculations, reconciliation or ambiguity. “A short question can be COMPLEX. A long factual list can be SIMPLE.”
  5. Generic data definitions: the application’s entities, relationships and metrics, without any test data.

A typical pair of decisions:

  • “For order-2, list the product, quantity and seller” → SIMPLE: “Direct bounded retrieval… No complex inference or policy.”
  • “Explain why its seller has not received the payout” → COMPLEX: “requires tracing allocations, payments, shipments, returns/refunds, and payout status, plus policy-based eligibility.”

JEV

JEV is TypeSafe’s decision model, which I wrote about in Right Model, Right Job. I gave it the same routing rubric, adapted to JEV’s typed API. The rubric text went in as the instructions of a typed SIMPLE/COMPLEX choice question, with the first line changed to “Choose SIMPLE or COMPLEX” instead of asking for a reason. That keeps the two close to a like-for-like comparison.

JEV returns a typed choice and a probability for each option:

{"tier": {"type": "choice", "choice": "COMPLEX",
"probabilities": {"SIMPLE": 0, "COMPLEX": 1}}}

That was its answer for “Why has Cedar not been paid for order-1?”

Both the LLM classifier and JEV ran in standalone mode rather than through LiteLLM’s built-in auto-routing. LiteLLM supports both an LLM classifier and JEV as built-in options, so the same setup can be configured inside the gateway.

OpenRouter Auto

OpenRouter’s Auto Router works differently from the other five. There is no prompt or rule set to write, and it doesn’t use my SIMPLE/COMPLEX labels. OpenRouter applies its own task classification, then ranks candidate models using its market data.

  • I restricted it to only my two models (allowed_models: openai/gpt-5-nano, openai/gpt-5).
  • I supplied the cost tier myself, trying low and medium.
  • Both settings selected nano for every case.

That is a finding about this restricted two-model configuration. It is not evidence that OpenRouter’s router performs poorly in general.

The Results

Here is how each approach did on the same 18 cases. Costs are replay estimates for all 18 cases at full list price, and they include the classifier’s own cost. The keyword and semantic rows include their heuristic fallback.

ApproachCases correct / partial / incorrectFully correctnano / GPT-5 picksEstimated costSaving vs always GPT-5
Always GPT-5 nano9 / 7 / 250.0%18 / 0$0.02394.9%
Always GPT-517 / 1 / 094.4%0 / 18$0.456—
Keyword rules15 / 2 / 183.3%8 / 10$0.24047.4%
Heuristic (default)9 / 7 / 250.0%17 / 1$0.03891.7%
Semantic similarity13 / 4 / 172.2%13 / 5$0.13869.8%
LLM classifier16 / 1 / 188.9%9 / 9$0.24945.3%
JEV16 / 1 / 188.9%8 / 10$0.26641.7%
OpenRouter Auto (low and medium)9 / 7 / 250.0%18 / 0$0.02394.9%
Scatter plot titled "18-case saved-answer replay: case-level correctness vs estimated cost". Always GPT-5 is top right at 94.4% and $0.456. The LLM classifier and JEV sit at 88.9% around $0.25 to $0.27, keyword rules with heuristic fallback at 83.3% and $0.240, semantic similarity with heuristic fallback at 72.2% and $0.138. The default heuristic, always nano and OpenRouter Auto all sit at 50% near $0.02 to $0.04.
On these 18 cases, classifiers written for the domain came closest to GPT-5’s case-level correctness at about half its estimated cost.

On this tested set:

  • The domain-aware classifiers did better than the generic ones. The LLM classifier and JEV share one domain-aware routing rubric. They kept 16 of GPT-5’s 17 fully correct cases at 55–58% of its estimated cost. The default heuristic and the restricted OpenRouter configuration selected nano almost every time and scored the same as nano alone.
  • Two short keyword lists got surprisingly close: 83.3% against 88.9%, at similar cost and with no model call.
  • Separate the classifier’s cost from the route’s cost. JEV’s 18 classification calls cost less than the LLM classifier’s (about $0.0038 versus $0.0054), and each decision was faster (about 0.7 s versus 2.7 s). Its route still cost more overall, because JEV selected GPT-5 for 10 cases against 9. The extra one was a case nano had already answered correctly.
  • Semantic matching depends on its examples. 10 of 18 cases matched no example closely enough and fell back to the heuristic. In the separately versioned refinement, I added more examples: fallbacks dropped to 6, but correctness did not improve and cost went up.
  • Routing did not outperform fixed GPT-5 here. That isn’t a law. Routing can beat the best fixed model when different models succeed on different cases. In this pair, GPT-5 was also right on every case where nano was right, so there was nothing for routing to gain over GPT-5 alone.

Right vs Wrong: Three Examples

Averages hide the interesting part, so here are three cases in detail.

1. A short question isn’t a simple question. “Why has Cedar not been paid for order-1?” is only eight words, and the default heuristic scored it as simple and sent it to nano. The real answer needs the payout records: the seller’s payout had failed because of a provider timeout. Nano’s answer did not mention the timeout and was graded partial. GPT-5 identified the recorded provider timeout and was correct. The keyword rules (“why”), the LLM classifier and JEV all sent this question to GPT-5. Generic signals measure how a question looks; domain-aware ones try to measure what answering it takes.

2. Two identical questions, routed differently. Two questions differed only in the order ID: “Explain the refund history for ord-6f3604a5fe.” and “Explain the refund history for ord-aaa1768930.” In the refined semantic version, the first scored 0.529 and the second 0.493 against the closest example, on opposite sides of the 0.5 threshold.

  • The first went to GPT-5, although nano’s answer to it was already fully correct. That was wasted cost.
  • The second fell back to the heuristic and went to nano, whose answer was partial while GPT-5’s was correct. That was lost accuracy.

It was wrong in both directions, because of an order ID. Similarity thresholds can be sensitive to details that don’t change the task.

3. A keyword list is only as good as its coverage. “After returning one Canvas Weekender, what remains payable to Cedar?” is a money calculation: take the seller’s share of the sale, then adjust for the refunded item and the fee. Its wording (“returning”, “payable”) wasn’t in either keyword list, so it fell back to the heuristic and went to nano. Nano’s answer was partial; GPT-5 computed the correct amount, ₹2,249.10. Rules are cheap and transparent, but this gap led to a routing mistake. A fallback can sometimes still choose correctly; here it didn’t.

Lessons Learned

Domain-aware configurations did better on this set. Two keyword lists reached 83.3%, versus 88.9% for the LLM classifier and JEV. That makes simple rules a useful baseline, with their coverage gaps tested explicitly.

The LLM classifier and JEV agreed on 17 of 18 choices. JEV was faster and cheaper per decision, but agreement does not establish which part of the setup caused the benefit. Separate classifier cost from selected-model cost: a cheap classifier can still produce an expensive route by choosing GPT-5 more often.

A Checklist for Choosing a Classifier

  1. Shortlist models on benchmarks, domain fit, cost and latency, from approved providers that support your agent’s needs.
  2. Run evals on your golden set to narrow the shortlist. Then fix the pool and pin the versions.
  3. Weigh cost savings against accuracy. If the cheaper model meets your accuracy bar, use it. If only the stronger one does and its cost is acceptable, use it. Route only when routing saves meaningful cost while keeping accuracy within your bar.
  4. Design the classification. Choose labels that matter for your traffic and map them to models. Start with keyword rules, then try semantic or LLM-based classifiers built from real user journeys, not from test questions. Compare them on the same golden set, and look at both classifier cost and route cost.

Part 2 adds the remaining steps: choosing the scope on real conversations, and validating before you deploy.

Limits and What’s Next

This was a small study, and I would not treat it as a benchmark:

  • Sample size: 18 cases (20 answer turns per model), one run each.
  • Grading: the earlier calibrated judge, existing accepted grades, and an LLM reviewer with my adjudication of borderline cases.
  • Tuning: the rules, examples and prompts were written by someone who knew the test cases.
  • Costs: replay estimates at full list price.

The point is the process, not the percentages. Next, I want to:

  • test on fresh held-out cases, with repeated runs
  • add a middle-tier model
  • try routers I haven’t tested yet: Not Diamond, which learns from your graded data; Aurelio’s semantic-router; RouteLLM; Arch-Router; and the cloud platforms’ own model routers

Acknowledgements

This post was written with the help of Claude Opus 5.5 in Claude Code, which also helped me research the routing landscape and plan the experiments. The application and experiments were built and run with GPT-6 Astra, which also reviewed answers against the expected results. Claude Opus 5.5 served as the calibrated LLM judge in the earlier model comparison.

References

This Series

Earlier Posts

Routing Tools

For Future Work