TL;DR: OpenAI’s Agents API (POST /v1/agents/sessions, public beta since September 10, 2026) turns “agent” from a while-loop you own into a durable, resumable object OpenAI hosts for you — model, instructions, tools, and a sandbox, addressed by one session_id. You stop rebuilding job queues and container orchestration for every agent project and start managing session state instead. This post covers what actually changed, the real pricing, where it breaks in practice, and a full worked example you can run today.


The loop you’ve probably already written by hand

If you’ve shipped anything agentic on OpenAI’s models, you’ve written this loop, or something close to it: call the model, check if it asked for a tool, run the tool, stuff the result back into the context, call the model again, repeat until it stops asking for tools. Then you wrapped it in a job queue so it could survive longer than one HTTP request. Then you built a container orchestrator so tool calls that run code had somewhere safe to run. Then you added a state machine so a crashed worker could resume a task instead of losing it.

None of that is agent logic. It’s plumbing every team building agents ends up writing, independently, slightly differently — and it’s exactly the layer OpenAI decided to sell as a managed service on September 10, 2026: the Agents API. It’s not a new model, and it’s not a replacement for the Agents SDK you might already be using — it’s a hosted version of the plumbing.

What it actually is

A session is the unit that replaces your loop. It bundles four things:

  • Agent — the model, its instructions, its tools, and any MCP servers it’s allowed to call
  • Environment — where the agent’s tool calls actually execute: an OpenAI-hosted sandbox, or one you bring yourself (E2B, Modal, DigitalOcean, Vercel are supported)
  • Session — the durable, resumable object doing the work, addressed by a session_id that outlives your client connection
  • Events/Items — the timestamped record of everything that happened, readable back at any time

This sits at a different layer from tools you’ve likely already used, and mixing them up is the single most common source of confusion in this space:

LayerWhat it isWho owns state
Responses APIOne model call, one responseNobody — it’s stateless
Agents SDKA Python/TypeScript framework you run: Agent, Runner, handoff()You, in your own process
Agents APIA managed session service: POST /v1/agents/sessionsOpenAI, durably, server-side

The Agents SDK didn’t get replaced — it got a sibling. If your server already owns tool execution and approval logic, the SDK still gives you tighter, typed, in-process control. The Agents API exists for the specific case where you’d rather not run that infrastructure yourself.

Why this is worth your attention right now

Long-running agent work fundamentally doesn’t fit the request/response model most backends are built around. “Refactor this repo, run the test suite, fix what’s red” can take ten minutes, needs to keep going if your server restarts, and needs a real filesystem and shell to run in — not a sandboxed eval(). Before September 2026, you built all of that yourself. Now it’s three HTTP calls:

POST /v1/agents/sessions               start a session
POST /v1/agents/sessions/{id}/events   steer it mid-turn, or continue it
GET  /v1/agents/sessions/{id}/items    read back what happened, paginated

And critically, a session is a real object with a lifecycle independent of your process. Your laptop can sleep, your Lambda can time out, your pod can get rescheduled — the session keeps running server-side, and you reconnect to it with the same session_id and pick up exactly where it left off. That’s the actual unlock. Everything else is convenience on top of it.

A turn is one cycle of work. Message an idle session and you start a new turn. Message a session that’s already mid-turn and you steer it — you redirect the current work without waiting for it to finish and re-prompting from scratch. That distinction matters more than it looks: steering keeps the model’s original context and reasoning intact, while “wait, then correct” throws it away and starts cold.

Worked example: a repo task, start to finish

Here’s the whole lifecycle, not just the create call everyone quotes.

Credentials and scopes

Your key needs api.agents.read and api.agents.write for session operations, plus api.responses.write for the model calls underneath. Every request also needs a beta header — this is still a beta API, and OpenAI will change it.

export OPENAI_API_KEY="your-api-key"

Start the session and stream it

from openai import OpenAI

with OpenAI() as client:
    with client.beta.agents.sessions.create(
        agent={
            "model": "gpt-6-astra",
            "instructions": "Write clean code, run it, and report the actual output.",
        },
        environment={"type": "openai_hosted"},
        input="Create tree.py, a Python script that prints a readable tree "
              "of the files in the current directory. Run it and show me the output.",
        stream=True,
    ) as events:
        for event in events:
            print(event.to_json(indent=None), flush=True)

Watch specifically for agent.session.turn.completed. OpenAI’s own docs are blunt about this: an idle session does not mean the turn succeeded. Silence is not success — check the event type explicitly, every time.

It’s halfway done and you want to redirect it — steer, don’t restart

curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/events \
  -H "OpenAI-Beta: agents=v1" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"events":[{"type":"agent.session.input.message","content":"Also add type hints."}]}'

It’s gone sideways — cancel the turn without losing the session

curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/events \
  -H "OpenAI-Beta: agents=v1" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"events":[{"type":"agent.session.input.cancel"}]}'

Cancelling stops the active turn and keeps everything before it — the session, its history, and its context are all still there for the next turn.

Audit what actually happened

curl https://api.openai.com/v1/agents/sessions/$SESSION_ID/items \
  -H "OpenAI-Beta: agents=v1" \
  -H "Authorization: Bearer $OPENAI_API_KEY"

This is your paginated, durable trail of every message and tool-call result — the thing you’d have built a logging pipeline for, given to you as a GET.

What it actually costs

There’s no separate “agent fee.” You pay the model at its normal API rate, plus standard rates for any hosted tools, plus container time for the sandbox — billed per minute, with a 5-minute minimum, in 20-minute pricing bands by memory size:

Sandbox memoryPrice per 20-minute session
1 GB$0.03
4 GB$0.12
16 GB$0.48
64 GB$1.92

Web search, when the agent uses it, is billed separately at $10 per 1,000 calls (plus the search content’s own token cost).

! Runaway sessions bill by the minute

None of these numbers are large in isolation — the actual risk is an agent stuck in a retry loop with a container open, or a demo you forgot to tear down. A developer flagged exactly this in OpenAI’s own launch thread: calculate the ceiling before you spin containers up, because the meter runs by the minute, not by the outcome.

Where teams are actually pointing this

OpenAI’s cookbook examples for the Agents API/SDK cluster around a specific shape of problem: long, semi-autonomous work with a real filesystem or shell underneath. SRE incident-response agents. Legacy codebase migrations run inside sandboxes. Prior-authorization review in healthcare. AML analysis with a security-evaluation layer on top. Database change-impact analysis that generates and checks SQL before anyone runs it. The common thread isn’t “chatbot” — it’s “task that would otherwise sit in someone’s queue for an hour.”

Common pitfalls

! Treating idle as succeeded

Worth repeating because it’s the single most common bug report: check for turn.completed vs. turn.failed explicitly. Don’t infer anything from the absence of activity.

! Forgetting the beta header

Every call needs OpenAI-Beta: agents=v1. Without it, the error you get reads like a permissions problem, not a missing-header problem — expect to lose ten minutes to this at least once.

⚠ Assuming this gets you Zero Data Retention

It doesn’t, even with a self-hosted sandbox, and data residency is US-only. If ZDR is a hard compliance requirement, build on the Agents SDK against your own infrastructure instead — this API is the wrong layer for you, full stop.

! Reaching for the Agents API when the SDK already fits

If your server already owns tool execution, approvals, and state — exactly what the SDK is for — layering the Agents API on top just gives you two orchestration layers arguing with each other. Pick one owner of durability: OpenAI, or you.

! Polling instead of steering

Waiting for a turn to finish, then re-prompting with a correction as a brand-new turn, is slower and throws away context framing the model already built. Steer mid-turn instead.

Which one should you actually use?

Reach for the Agents API when…Reach for the Agents SDK when…Reach for the plain Responses API when…
You want OpenAI to host the sandbox and own session durabilityYour server already owns deployment, tools, and stateIt’s one call, no persistence — classification, extraction, a single tool call
Tasks run long and need to survive client disconnectsYou need typed, in-process control over handoffs and guardrailsYou don’t need a sandbox, a session, or a session_id at all
US-only data residency is acceptableYou need Zero Data RetentionBringing session infrastructure here is pure overhead
★ Key takeaways
  • The Agents API is a managed session service (/v1/agents/sessions), public beta since September 10, 2026 — a sibling to the Agents SDK, not a replacement for it.
  • A session = Agent (model + instructions + tools) + Environment (sandbox) + durable state, addressed by one session_id that survives your client disconnecting.
  • A turn is one work cycle. Message an idle session to start one; message an active one to steer it — steering beats poll-then-reprompt.
  • Never infer success from silence. Check explicitly for agent.session.turn.completed.
  • Pricing is model rate + tool rate + sandbox time — no separate “agent fee,” but runaway sessions bill by the minute regardless of outcome.
  • It’s US-only and doesn’t support Zero Data Retention, even self-hosted — know this before you architect around it.
  • The decision isn’t “which is better,” it’s “who should own durability and sandboxing — OpenAI, or you.”

Further reading