AI / A CONCEPT NOTE

LoRA & QLoRA

fine-tuning models on consumer hardware

~65 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Low-Rank Adaptation. Freezes the base model weights and injects small, trainable adapters (rank decomposition matrices) into attention layers, reducing training memory by 90%+.

02 / FOLLOW THE MECHANISM

How adapter training flows

  1. Base model freeze

    locks all pre-trained weights to prevent costly changes.

  2. Adapter injection

    inserts small parallel layers (LoRA weight matrices) into neural layers.

  3. Forward / Backward

    routes inputs through both; updates only the small adapter weights.

  4. Weight merging

    adds adapter parameters back into base model for fast zero-overhead serving.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

merge adapter weights

peft-convert --base llama-3 --adapter ./lora

Explore command anatomy in the CLI lab