AI / A CONCEPT NOTE
RAG
grounding LLM answers in your own data
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Retrieval-Augmented Generation (RAG) adds a retrieval step before the LLM generates an answer. You embed a user's question into a vector, search a database of document chunks for the most relevant content, inject those chunks into the prompt as context, and then the LLM answers based on that retrieved information — not just its training data.
02 / FOLLOW THE MECHANISM
How a RAG query flows
User
asks a question: 'What's our incident response SLA?'
Embedding model
converts the question into a vector (numeric embedding) that captures its semantic meaning.
Vector database
finds the top-k document chunks whose embeddings are nearest to the question's embedding (cosine similarity).
Context builder
stuffs the retrieved chunks into the LLM prompt as context, above the user's question.
LLM
generates an answer using only the provided context — it doesn't guess from its training data.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
generate an embedding for a query
curl -X POST -d '{"input":"What is RAG?"}' https://api.openai.com/v1/embeddingsquery for nearest vectors in pgvector
pgvector 'SELECT * FROM docs ORDER BY embedding <=> $1 LIMIT 5'05 / CHECK YOURSELF
Could you explain RAG to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringVector Databasesstoring and searching by semantic similarity, not exact keywords