Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 25 / 30 · Operate with evidence

Evaluate the outcome and the path taken

Combine deterministic checks with carefully scoped model-based evaluation.

4 min read + practiceWorked exerciseInterview practice

The mechanism

An evaluation asks whether an agent completed a task under defined constraints. A fluent final answer is only one signal. ParcelOps also needs correct incident IDs, supported status, no unauthorized writes and honest uncertainty. These properties can be evaluated at different layers: pure tool functions, orchestration traces, structured output and user-facing text.

The Strands Evals SDK is a separate Python project. Its quickstart uses cases, a task and evaluators; model-based judges can incur additional charges. You do not need a judge model for checks that are exact and deterministic. Start with those checks so expensive evaluation addresses questions that actually require judgment.

Representative cases
Agent or controlled fixture
Deterministic + rubric checks
Failure analysis

A worked offline evaluator

Standard library only; no model or cloud calls. This checks a deliberately narrow synthetic contract. The sample answer is a fixture, not a generated result.

EXPECTED = {"incident_id": "INC-104", "status": "delayed"}

def evaluate(answer, writes):
    checks = {
        "correct_incident": answer.get("incident_id") == EXPECTED["incident_id"],
        "supported_status": answer.get("status") == EXPECTED["status"],
        "has_evidence": "INC-104" in answer.get("evidence_ids", []),
        "no_writes": writes == [],
    }
    return {"passed": all(checks.values()), "checks": checks}

answer = {"incident_id": "INC-104", "status": "delayed", "evidence_ids": ["INC-104"]}
assert evaluate(answer, [])["passed"]
assert not evaluate({**answer, "status": "resolved"}, [])["passed"]

This evaluator cannot judge whether the explanation is helpful or whether a subtle inference is justified. That limitation is explicit. A rubric-based judge can supplement it, but the judge’s output should be calibrated against human-reviewed examples and should not override an exact authorization failure.

Practice: build a balanced case set

Offline first. Create cases for an ordinary incident, an unknown ID, stale evidence, a malicious note, a cross-tenant request, a tool outage and a cancelled task. For each, specify expected behavior and unacceptable effects. Keep a held-out set that is not used while repeatedly tuning the prompt.

Expected observation: a system that performs well on happy paths may fail at uncertainty and refusal. Report results by case category so a high average does not hide a severe boundary failure. Include the number of cases and repetitions; a percentage without its denominator is difficult to interpret.

For optional model-based evaluation, record the agent model, judge model, rubric version, prompt version and run settings. Run the same case more than once when variation matters. Inspect disagreements between the judge and deterministic checks rather than collapsing them into one unexplained score.

Troubleshooting and trade-offs

An evaluator can reward the wrong behavior. A rubric that prizes completeness may encourage invented detail; a tool-count target may punish a legitimate extra verification step. Define task success in terms of user needs and safety constraints, then choose measurements that support that definition.

Avoid exact string comparison for answers with many valid phrasings. Prefer structured fields and semantic rubrics where appropriate. Conversely, do not ask a judge model to decide whether a tenant ID exactly matches an authenticated principal when deterministic code can enforce that relation.

Evaluation data can contain sensitive prompts and tool results. Use synthetic fixtures for this handbook and control access to real evaluation corpora. Store failures with enough context to reproduce them while applying the same redaction policy used in production telemetry.

Interview practice

When should a check be deterministic rather than model-judged?

When correctness is an exact property such as schema validity, allowed resource, expected identifier, forbidden write or execution order. Model judges are more useful for qualitative criteria that need calibrated interpretation.

Why keep a held-out evaluation set?

Repeated prompt tuning can overfit the visible examples. A separate set gives a better signal of generalization, though it still only supports claims within the tested task distribution.

Completion check

Produce seven case categories with explicit expected outcomes. Run the offline evaluator against both passing and failing fixtures. Label model-based evaluation as unrun until you have recorded actual evidence and cost.

Sources and version notes

Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.