AI / A CONCEPT NOTE

Model Serving

putting a model behind an endpoint

~70 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Model serving means deploying a trained model behind an HTTP/gRPC endpoint (like vLLM, Triton) that dynamically batches queries to optimize graphics card (GPU) utility.

02 / FOLLOW THE MECHANISM

How dynamic batching works

  1. Client queries

    multiple users send text prompts to the serving endpoint simultaneously.

  2. Batch engine

    vLLM groups prompts into a single matrix block (continuous batching).

  3. GPU execution

    evaluates batched matrix queries in a single hardware cycle.

  4. Stream response

    streams tokens back to users asynchronously as they complete.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

start an OpenAI-compatible serving host locally with vLLM

python -m vllm.entrypoints.openai.api_server --model facebook/opt-125m

Explore command anatomy in the CLI lab