AI / A CONCEPT NOTE

Evals

testing something non-deterministic

~65 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Evaluating LLM outputs by comparing model responses against test benchmarks (accuracy, toxicity, safety) using deterministic assertions or evaluator models.

02 / FOLLOW THE MECHANISM

How an eval runs

  1. Prompt test set

    defines 100 benchmark prompts alongside target ground-truth answers.

  2. Model execution

    runs prompts through the target candidate LLM to collect outputs.

  3. Scoring

    evaluator model or script scores answers: semantic similarity, keyword checking, etc.

  4. Performance report

    outputs accuracy metrics: 94% passed, 6% failed.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

run a prompt engineering evaluation matrix test locally

promptfoo eval

Explore command anatomy in the CLI lab