Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 07 / 30 · Design reliable tools

Design tool contracts around decisions

Use precise inputs, bounded outputs and error categories that support the next safe action.

4 min read + practiceWorked exerciseInterview practice

The mechanism

A tool contract is a promise about input, behavior and output. It should make the next decision easier for both the agent and the application. A generic string response such as “something went wrong” hides whether the request was invalid, forbidden, temporarily unavailable or already completed. Those cases require different responses.

For ParcelOps, a read operation returns an incident snapshot with its revision and observation time. A proposed write carries an expected revision and a request identifier. Separating these operations prevents a tool named “lookup” from unexpectedly changing state. It also gives reviewers a clear place to inspect side effects.

Intent-specific input
Validation + permission
Bounded operation
Typed outcome + reference

A worked result envelope

This is an application contract, not a built-in Strands response type. Your tool can return such a dictionary as domain data. Keep the SDK’s own tool-result protocol separate from your business fields.

{
  "outcome": "found",
  "incident_id": "INC-104",
  "revision": 7,
  "observed_at": "2026-10-06T10:00:00Z",
  "record": {"status": "delayed", "cause": "carrier scan missing"},
  "source_ref": "fixture:incidents/INC-104@7"
}

This synthetic timestamp is an example value, not a recorded run. The envelope answers four questions: what happened, which resource was observed, how fresh the observation is and where it came from. A stable source reference helps a final answer cite evidence without repeating an entire private record.

Define errors equally carefully. invalid_input invites correction; not_available may invite a bounded retry; forbidden must not trigger credential guessing; conflict requires rereading the current revision. Whether you expose “forbidden” to a user depends on enumeration risk. Internal audit semantics can be richer than the public response.

Practice: reduce an overpowered tool

Offline. Start with this unsafe interface: query_system(command: str). It accepts arbitrary operations and returns unrestricted output. Replace it on paper with read_incident(incident_id), propose_note(incident_id, text) and a separately authorized apply_note(proposal_id).

For each, list allowed resources, maximum input size, maximum result size, timeout, side effects and error categories. Expected observation: the smaller functions require more deliberate design but make privilege review and testing much more concrete. The “apply” operation can be absent entirely in a read-only deployment.

Now add a record with a 5 MB carrier note. Decide whether to truncate, paginate or return a reference. Preserve enough metadata to show that content was omitted. Silent truncation can make an answer falsely appear complete, so the result should state its limits.

Troubleshooting and trade-offs

Overly permissive schemas often arise from convenience: accepting arbitrary URLs, paths or SQL saves implementation effort while moving risk into the model. Prefer identifiers that resolve through a trusted service. If users need flexible search, expose bounded filters instead of an unrestricted backend query language.

Overly strict schemas can also be harmful when they reject legitimate domain values or make retries futile. Keep validation aligned with the actual resource contract, and return a useful correction message. Do not expose raw stack traces as “helpful detail”; store them in restricted diagnostics and provide a correlation reference.

Interview practice

Why separate propose and apply operations?

They have different authority and evidence requirements. A proposal can be reviewed and hashed before a trusted executor validates approval, current state and policy. It also supports read-only deployments.

What information makes a tool error actionable?

A stable category, a safe explanation, whether retry is appropriate, and a correlation reference. Include only information the caller may see, and keep sensitive diagnostics out of model-visible content.

Completion check

Write a contract for one read and one proposed write. Explain which fields come from trusted application state and which may be supplied by the model. Include a large-output and a stale-revision test.

Sources and version notes

Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.