AI / A CONCEPT NOTE
Quantization
shrinking models to fit
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Quantization compresses model parameter weights by converting high-precision numbers (FP32/FP16) to lower-precision formats (INT8/INT4), shrinking file sizes.
02 / FOLLOW THE MECHANISM
How quantization compresses weights
Base weights
large model weights are stored in floating point formats (16-bit).
Quantization run
compiles floating point scales to integer scales (e.g. 4-bit INT4).
Shrink write
writes compressed checkpoints, shrinking files by up to 75%.
Execution
model runs with low VRAM footprint, losing minimal intelligence accuracy.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
check quantization precision of a local model
ollama show llama3 | grep -i quantization05 / CHECK YOURSELF
Could you explain Quantization to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringTemperature & Samplingwhy answers vary