AI / A CONCEPT NOTE

Attention

how models calculate context-sensitive word relationships

~75 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Self-attention computes Query-Key-Value (Q, K, V) dot products allowing tokens to weigh and dynamically route context from other tokens across the sequence.

02 / FOLLOW THE MECHANISM

How attention scores are computed

  1. Input Token

    is projected into Query, Key, and Value vectors.

  2. Dot Product (Q · K)

    calculates similarity scores between token pairs.

  3. Softmax Scaling

    normalizes similarity into attention weight percentages.

  4. Weighted Value Sum

    aggregates context-rich representations for downstream layers.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

compute QK^T score matrix

python -c import torch; q=torch.randn(1,8,64); k=torch.randn(1,8,64); print((q @ k.transpose(-2,-1)).shape)

Explore command anatomy in the CLI lab