Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 13 / 30 · Capabilities

Compare agent integrations fairly

Apply the same task contract to Claude, Gemini and OpenCode while respecting their different setup paths.

4 min readWorked exerciseInterview practice

What you will build

A comparison protocol, not a leaderboard. Use chapter 12’s parser task so different integrations are assessed against the same input and acceptance rules.

The mechanism at a glance
  1. One input revision
  2. Separate agent environments
  3. Same acceptance contract
  4. Compare artifacts and failures

Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.

Read the mechanism

The microVM provides an execution boundary; each agent still has its own authentication, startup arguments, configuration and model selection. Replacing an agent name in a command does not prove equivalent behavior.

Docker documents Claude API-key and subscription authentication paths. OpenCode supports several configured providers and has no implicit launch flags in its integration. Gemini’s current page contains conflicting statements about interactive OAuth: this edition relies on its documented Google API-key path and treats interactive OAuth as requiring installed-version verification.

Do not compare “agents” while changing the model, tool access, task context and timeout at the same time. Even cache warmth or a previous failed attempt left in the workspace can change the outcome.

Worked example · a controlled comparison card

For each selected integration, fill out the following card before launching:

Input commit:
Agent / installed version:
Template / immutable identifier:
Provider / model:
Workspace mode:
Skills and MCP exposure:
Permitted network destinations:
CPU / memory:
Time and model-spend limit:
Exact prompt:
Acceptance test:
Retries allowed:
Export and cleanup method:

Create a separate clone from the same clean main checkout for each agent you actually intend to test:

sbx create --clone --name handbook-claude --skills off claude .
sbx create --clone --name handbook-gemini --skills off gemini .
sbx create --clone --name handbook-opencode --skills off opencode .

These are optional resource-creating commands. Run only the integrations for which you have approved credentials and a spending limit. Attach to one named sandbox at a time and use the same task contract. Keep the provider’s secret setup on the host.

Record failure behavior alongside successful patches. Did the agent stop on a policy denial, request a broader permission, modify tests, or change dependencies? Those are operational findings, even when no final patch was produced.

Expected observations

You should obtain comparable evidence packets, each clearly tied to an input SHA and configuration. Outcomes may differ. One successful run does not establish a general accuracy ranking, and an authentication failure is not evidence about reasoning quality.

If the exercise remains document-only, mark every result “not run.” A complete protocol is useful work; invented scores are not.

Troubleshooting

When an argument behaves differently, inspect the agent integration’s current startup rules. Flags after -- are agent arguments, not necessarily sandbox flags. When provider choice is ambiguous, inspect the documented configuration before making a billable request.

Gemini OAuth uncertainty should remain visible in the record. Do not silently claim both authentication flows were validated.

Interview practice

Why is one run per model a weak benchmark?

Stochastic behavior, cache state and task selection can dominate a single observation. Use repeated comparable trials and report failures and variance.

What stays constant when changing the agent?

The source revision, task, acceptance tests, allowed tools, resource and time budget, and review method. Record the settings you cannot hold constant.

Completion check

Produce one fully specified comparison card and identify at least three confounders. Leave performance and cost fields unfilled until observed.

Sources and version notes

Checked 6 October 2026; current baseline: sbx v0.46.0. Claude Code integration · Gemini integration · OpenCode integration

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.