The mechanism
A tool has two audiences. The model reads its name, description and argument schema to decide whether it is useful. Your application executes its implementation. Python’s @tool decorator connects a function to the SDK using its type hints and documentation, but it does not remove the need for ordinary unit tests and explicit validation.
We will use an in-memory fixture with one record. Keeping the lookup pure makes its behavior easy to reason about. The function does not access the network, accept a filesystem path or expose credentials. Narrow tools are easier to authorize and evaluate than a generic “run any command” tool.
A worked implementation
Local SDK example, no model call in this block. The plain function can be tested before it is wrapped. The returned dictionary is application data, not a claim that a live incident service was queried.
from strands import tool
INCIDENTS = {
"INC-104": {"status": "delayed", "cause": "carrier scan missing"}
}
def lookup_record(incident_id: str) -> dict:
if not incident_id.startswith("INC-") or not incident_id[4:].isdigit():
raise ValueError("Expected an incident ID such as INC-104")
record = INCIDENTS.get(incident_id)
return {"incident_id": incident_id, "found": record is not None,
"record": dict(record) if record else None}
@tool
def lookup_incident(incident_id: str) -> dict:
"""Read one synthetic incident. Does not modify incident state.
Args:
incident_id: Exact identifier, for example INC-104.
"""
return lookup_record(incident_id)
assert lookup_record("INC-104")["found"] is True
assert lookup_record("INC-999")["found"] is False
Pass lookup_incident in Agent(tools=[lookup_incident], model=model) when you intentionally enable inference. That model object is the explicitly configured one from chapter 3. The assertions above exercise application lookup behavior; they do not prove that a model will select the tool at the right time.
Returning a copied record prevents a caller from mutating the fixture through the returned dictionary. This is a small example of a broader principle: make the authority and mutation behavior of every interface obvious. In a real service, enforce tenant and field permissions at the data boundary, not by trusting the incident ID format.
Practice: test selection separately from execution
Offline first. Add tests for a malformed ID, a missing record and an existing record. Then describe three prompts that should cause a lookup and three that should not. “Explain INC-104” needs the fixture; “What is an incident ID?” may not. A model-selection evaluation measures whether the agent chooses appropriately, while the unit tests measure what the tool does once called.
Expected observation: a tool can pass every unit test while the agent never uses it. If that happens in the optional live track, improve the tool’s description and task instructions, then rerun the same selection cases. Do not broaden its authority merely to make selection easier.
Troubleshooting and trade-offs
A vague tool name such as helper gives the model little information. A description that advertises capabilities the code lacks creates a misleading contract. If a result is huge, return only relevant fields and a stable reference to the source. If the tool raises an error, distinguish invalid input from unavailable infrastructure so retry logic can make an informed decision.
Avoid including secrets or internal exception traces in tool errors. A model-visible error should help the agent choose the next safe step without revealing implementation details. Preserve deeper diagnostics in a restricted application log with a correlation ID.
Interview practice
Why test the plain function independently of the agent?
It separates deterministic business logic from probabilistic tool selection. Failures become easier to localize, and most input, permission and mutation cases can be checked without paying for inference.
What makes a useful tool description?
A precise purpose, argument meaning, relevant constraints, side effects and result semantics. It should help selection while accurately reflecting the implementation; it is not a replacement for enforcement.
Completion check
Demonstrate existing, missing and malformed inputs against the pure function. Explain the separate evidence needed to show correct model selection and correct tool execution.
Sources and version notes
Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.