AI / A CONCEPT NOTE
Attention
how models calculate context-sensitive word relationships
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Self-attention computes Query-Key-Value (Q, K, V) dot products allowing tokens to weigh and dynamically route context from other tokens across the sequence.
02 / FOLLOW THE MECHANISM
How attention scores are computed
Input Token
is projected into Query, Key, and Value vectors.
Dot Product (Q · K)
calculates similarity scores between token pairs.
Softmax Scaling
normalizes similarity into attention weight percentages.
Weighted Value Sum
aggregates context-rich representations for downstream layers.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
compute QK^T score matrix
python -c import torch; q=torch.randn(1,8,64); k=torch.randn(1,8,64); print((q @ k.transpose(-2,-1)).shape)05 / CHECK YOURSELF
Could you explain Attention to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringTransformersthe foundational neural network architecture of modern LLMs