TL;DR: OpenAI’s Agents API (
POST /v1/agents/sessions, public beta since September 10, 2026) turns “agent” from a while-loop you own into a durable, resumable object OpenAI hosts for you — model, instructions, tools, and a sandbox, addressed by onesession_id. You stop rebuilding job queues and container orchestration for every agent project and start managing session state instead. This post covers what actually changed, the real pricing, where it breaks in practice, and a full worked example you can run today.
The loop you’ve probably already written by hand
If you’ve shipped anything agentic on OpenAI’s models, you’ve written this loop, or something close to it: call the model, check if it asked for a tool, run the tool, stuff the result back into the context, call the model again, repeat until it stops asking for tools. Then you wrapped it in a job queue so it could survive longer than one HTTP request. Then you built a container orchestrator so tool calls that run code had somewhere safe to run. Then you added a state machine so a crashed worker could resume a task instead of losing it.
None of that is agent logic. It’s plumbing every team building agents ends up writing, independently, slightly differently — and it’s exactly the layer OpenAI decided to sell as a managed service on September 10, 2026: the Agents API. It’s not a new model, and it’s not a replacement for the Agents SDK you might already be using — it’s a hosted version of the plumbing.
What it actually is
A session is the unit that replaces your loop. It bundles four things:
- Agent — the model, its instructions, its tools, and any MCP servers it’s allowed to call
- Environment — where the agent’s tool calls actually execute: an OpenAI-hosted sandbox, or one you bring yourself (E2B, Modal, DigitalOcean, Vercel are supported)
- Session — the durable, resumable object doing the work, addressed by a
session_idthat outlives your client connection - Events/Items — the timestamped record of everything that happened, readable back at any time
This sits at a different layer from tools you’ve likely already used, and mixing them up is the single most common source of confusion in this space:
| Layer | What it is | Who owns state |
|---|---|---|
| Responses API | One model call, one response | Nobody — it’s stateless |
| Agents SDK | A Python/TypeScript framework you run: Agent, Runner, handoff() | You, in your own process |
| Agents API | A managed session service: POST /v1/agents/sessions | OpenAI, durably, server-side |
The Agents SDK didn’t get replaced — it got a sibling. If your server already owns tool execution and approval logic, the SDK still gives you tighter, typed, in-process control. The Agents API exists for the specific case where you’d rather not run that infrastructure yourself.
Why this is worth your attention right now
Long-running agent work fundamentally doesn’t fit the request/response model most backends are built around. “Refactor this repo, run the test suite, fix what’s red” can take ten minutes, needs to keep going if your server restarts, and needs a real filesystem and shell to run in — not a sandboxed eval(). Before September 2026, you built all of that yourself. Now it’s three HTTP calls:
POST /v1/agents/sessions start a session
POST /v1/agents/sessions/{id}/events steer it mid-turn, or continue it
GET /v1/agents/sessions/{id}/items read back what happened, paginated
And critically, a session is a real object with a lifecycle independent of your process. Your laptop can sleep, your Lambda can time out, your pod can get rescheduled — the session keeps running server-side, and you reconnect to it with the same session_id and pick up exactly where it left off. That’s the actual unlock. Everything else is convenience on top of it.
A turn is one cycle of work. Message an idle session and you start a new turn. Message a session that’s already mid-turn and you steer it — you redirect the current work without waiting for it to finish and re-prompting from scratch. That distinction matters more than it looks: steering keeps the model’s original context and reasoning intact, while “wait, then correct” throws it away and starts cold.
Worked example: a repo task, start to finish
Here’s the whole lifecycle, not just the create call everyone quotes.
Credentials and scopes
Your key needs api.agents.read and api.agents.write for session operations, plus api.responses.write for the model calls underneath. Every request also needs a beta header — this is still a beta API, and OpenAI will change it.
export OPENAI_API_KEY="your-api-key"Start the session and stream it
from openai import OpenAI
with OpenAI() as client:
with client.beta.agents.sessions.create(
agent={
"model": "gpt-6-astra",
"instructions": "Write clean code, run it, and report the actual output.",
},
environment={"type": "openai_hosted"},
input="Create tree.py, a Python script that prints a readable tree "
"of the files in the current directory. Run it and show me the output.",
stream=True,
) as events:
for event in events:
print(event.to_json(indent=None), flush=True)Watch specifically for agent.session.turn.completed. OpenAI’s own docs are blunt about this: an idle session does not mean the turn succeeded. Silence is not success — check the event type explicitly, every time.
It’s halfway done and you want to redirect it — steer, don’t restart
curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/events \
-H "OpenAI-Beta: agents=v1" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"events":[{"type":"agent.session.input.message","content":"Also add type hints."}]}'It’s gone sideways — cancel the turn without losing the session
curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/events \
-H "OpenAI-Beta: agents=v1" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"events":[{"type":"agent.session.input.cancel"}]}'Cancelling stops the active turn and keeps everything before it — the session, its history, and its context are all still there for the next turn.
Audit what actually happened
curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/items \
-H "OpenAI-Beta: agents=v1" \
-H "Authorization: Bearer $OPENAI_API_KEY"This is your paginated, durable trail of every message and tool-call result — the thing you’d have built a logging pipeline for, given to you as a GET.
What it actually costs
There’s no separate “agent fee.” You pay the model at its normal API rate, plus standard rates for any hosted tools, plus container time for the sandbox — billed per minute, with a 5-minute minimum, in 20-minute pricing bands by memory size:
| Sandbox memory | Price per 20-minute session |
|---|---|
| 1 GB | $0.03 |
| 4 GB | $0.12 |
| 16 GB | $0.48 |
| 64 GB | $1.92 |
Web search, when the agent uses it, is billed separately at $10 per 1,000 calls (plus the search content’s own token cost).
None of these numbers are large in isolation — the actual risk is an agent stuck in a retry loop with a container open, or a demo you forgot to tear down. A developer flagged exactly this in OpenAI’s own launch thread: calculate the ceiling before you spin containers up, because the meter runs by the minute, not by the outcome.
Where teams are actually pointing this
OpenAI’s cookbook examples for the Agents API/SDK cluster around a specific shape of problem: long, semi-autonomous work with a real filesystem or shell underneath. SRE incident-response agents. Legacy codebase migrations run inside sandboxes. Prior-authorization review in healthcare. AML analysis with a security-evaluation layer on top. Database change-impact analysis that generates and checks SQL before anyone runs it. The common thread isn’t “chatbot” — it’s “task that would otherwise sit in someone’s queue for an hour.”
Common pitfalls
Worth repeating because it’s the single most common bug report: check for turn.completed vs. turn.failed explicitly. Don’t infer anything from the absence of activity.
Every call needs OpenAI-Beta: agents=v1. Without it, the error you get reads like a permissions problem, not a missing-header problem — expect to lose ten minutes to this at least once.
It doesn’t, even with a self-hosted sandbox, and data residency is US-only. If ZDR is a hard compliance requirement, build on the Agents SDK against your own infrastructure instead — this API is the wrong layer for you, full stop.
If your server already owns tool execution, approvals, and state — exactly what the SDK is for — layering the Agents API on top just gives you two orchestration layers arguing with each other. Pick one owner of durability: OpenAI, or you.
Waiting for a turn to finish, then re-prompting with a correction as a brand-new turn, is slower and throws away context framing the model already built. Steer mid-turn instead.
Which one should you actually use?
| Reach for the Agents API when… | Reach for the Agents SDK when… | Reach for the plain Responses API when… |
|---|---|---|
| You want OpenAI to host the sandbox and own session durability | Your server already owns deployment, tools, and state | It’s one call, no persistence — classification, extraction, a single tool call |
| Tasks run long and need to survive client disconnects | You need typed, in-process control over handoffs and guardrails | You don’t need a sandbox, a session, or a session_id at all |
| US-only data residency is acceptable | You need Zero Data Retention | Bringing session infrastructure here is pure overhead |
- The Agents API is a managed session service (
/v1/agents/sessions), public beta since September 10, 2026 — a sibling to the Agents SDK, not a replacement for it. - A session = Agent (model + instructions + tools) + Environment (sandbox) + durable state, addressed by one
session_idthat survives your client disconnecting. - A turn is one work cycle. Message an idle session to start one; message an active one to steer it — steering beats poll-then-reprompt.
- Never infer success from silence. Check explicitly for
agent.session.turn.completed. - Pricing is model rate + tool rate + sandbox time — no separate “agent fee,” but runaway sessions bill by the minute regardless of outcome.
- It’s US-only and doesn’t support Zero Data Retention, even self-hosted — know this before you architect around it.
- The decision isn’t “which is better,” it’s “who should own durability and sandboxing — OpenAI, or you.”
