AI / A CONCEPT NOTE

Model Routing

sending each prompt to the best model for the job

~75 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Model routing is the practice of dynamically selecting which LLM handles a given request based on criteria like complexity, cost, latency, or domain. A lightweight classifier (or simple heuristics) analyzes the incoming prompt and routes it — 'What time is it?' goes to a cheap, fast model; 'Debug this distributed system deadlock' goes to the most capable one. This optimizes spend and response quality simultaneously.

02 / FOLLOW THE MECHANISM

How a model router decides

  1. User prompt

    arrives at the routing layer: 'Explain the CAP theorem with examples.'

  2. Classifier

    evaluates complexity, token count, and domain tags — rates this as 'medium complexity, technical.'

  3. Routing rules

    match the classification to a model tier: simple → GPT-4o-mini, medium → Claude 3.5 Sonnet, hard → GPT-4o or Opus.

  4. Selected model

    receives the prompt and generates the response. The router logs which model was used, latency, and cost for optimization.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · INCOMPLETE SKETCH

use LiteLLM's router for multi-model routing

python -c "from litellm import Router; router = Router(model_list=[...]); router.completion(model='gpt-4o', messages=[...])"

The ellipsis omits required code or values. This sketch is not runnable as written.

Explore command anatomy in the CLI lab