AI / A CONCEPT NOTE
LoRA & QLoRA
fine-tuning models on consumer hardware
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Low-Rank Adaptation. Freezes the base model weights and injects small, trainable adapters (rank decomposition matrices) into attention layers, reducing training memory by 90%+.
02 / FOLLOW THE MECHANISM
How adapter training flows
Base model freeze
locks all pre-trained weights to prevent costly changes.
Adapter injection
inserts small parallel layers (LoRA weight matrices) into neural layers.
Forward / Backward
routes inputs through both; updates only the small adapter weights.
Weight merging
adds adapter parameters back into base model for fast zero-overhead serving.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
merge adapter weights
peft-convert --base llama-3 --adapter ./lora05 / CHECK YOURSELF
Could you explain LoRA & QLoRA to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringDSPycompiling declarative prompt pipelines