AI / A CONCEPT NOTE

Local LLM Inference

running open-source models on your developer machine

~75 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Running AI models entirely on your own hardware without sending data to third-party APIs. You download quantized model weights (like Llama 3) and run them locally using your system's VRAM/RAM.

02 / FOLLOW THE MECHANISM

How offline inference flows

  1. You

    run a command requesting a model like 'Llama 3'.

  2. CLI helper

    downloads the model parameters and loads them into your computer's graphics memory.

  3. Local Server

    starts a private API endpoint running on your localhost.

  4. Client

    sends prompts and receives responses offline with complete data privacy.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

pull and start a local chat session with Llama 3

ollama run llama3

EXAMPLE 02 · REFERENCE

inspect the default system prompt template

ollama show --system llama3

Explore command anatomy in the CLI lab