AI / A CONCEPT NOTE

Quantization

shrinking models to fit

~65 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Quantization compresses model parameter weights by converting high-precision numbers (FP32/FP16) to lower-precision formats (INT8/INT4), shrinking file sizes.

02 / FOLLOW THE MECHANISM

How quantization compresses weights

  1. Base weights

    large model weights are stored in floating point formats (16-bit).

  2. Quantization run

    compiles floating point scales to integer scales (e.g. 4-bit INT4).

  3. Shrink write

    writes compressed checkpoints, shrinking files by up to 75%.

  4. Execution

    model runs with low VRAM footprint, losing minimal intelligence accuracy.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

check quantization precision of a local model

ollama show llama3 | grep -i quantization

Explore command anatomy in the CLI lab