AI / A CONCEPT NOTE

RAG

grounding LLM answers in your own data

~80 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Retrieval-Augmented Generation (RAG) adds a retrieval step before the LLM generates an answer. You embed a user's question into a vector, search a database of document chunks for the most relevant content, inject those chunks into the prompt as context, and then the LLM answers based on that retrieved information — not just its training data.

02 / FOLLOW THE MECHANISM

How a RAG query flows

  1. User

    asks a question: 'What's our incident response SLA?'

  2. Embedding model

    converts the question into a vector (numeric embedding) that captures its semantic meaning.

  3. Vector database

    finds the top-k document chunks whose embeddings are nearest to the question's embedding (cosine similarity).

  4. Context builder

    stuffs the retrieved chunks into the LLM prompt as context, above the user's question.

  5. LLM

    generates an answer using only the provided context — it doesn't guess from its training data.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

generate an embedding for a query

curl -X POST -d '{"input":"What is RAG?"}' https://api.openai.com/v1/embeddings

EXAMPLE 02 · REFERENCE

query for nearest vectors in pgvector

pgvector 'SELECT * FROM docs ORDER BY embedding <=> $1 LIMIT 5'

Explore command anatomy in the CLI lab