AI / A CONCEPT NOTE
Evals
testing something non-deterministic
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Evaluating LLM outputs by comparing model responses against test benchmarks (accuracy, toxicity, safety) using deterministic assertions or evaluator models.
02 / FOLLOW THE MECHANISM
How an eval runs
Prompt test set
defines 100 benchmark prompts alongside target ground-truth answers.
Model execution
runs prompts through the target candidate LLM to collect outputs.
Scoring
evaluator model or script scores answers: semantic similarity, keyword checking, etc.
Performance report
outputs accuracy metrics: 94% passed, 6% failed.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
run a prompt engineering evaluation matrix test locally
promptfoo eval05 / CHECK YOURSELF
Could you explain Evals to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringLatent Spacehigh-dimensional space of meanings