OpenAI Agents API (Managed Agents) — architecture, lifecycle, streaming
Status: DOCUMENTED · BETA (OpenAI-Beta: agents=v1) · LIVE_VERIFIED (23 of 47 Agents/Vaults operations called successfully on 2026-09-18 with a standard project key — agents CRUD, sessions CRUD + events/items/turns/subagents/artifacts lists, environment-template CRUD, vault list; 4 more LIVE_DISCOVERED via 404 probes; the rest DOCUMENTED only).
Sources: Agents API overview · Architecture · Quickstart · Configuring agents · Sessions · Events and items · Manage sessions · Webhooks · Multi-agent · Functions · MCP · Observability · Tracing · Reference (beta/agents) · Streaming events reference · OpenAPI openapi-master.yaml (paths /agents/**, /vaults/**).
Last verified: 2026-09-18. Machine-readable twins: generated/fragments/endpoints/openai-agents.json, parameters/openai-agents.json, streaming-events/openai-agents.json, objects/openai-agents-objects.json.
The Agents API gives an application access to the Codex harness run by OpenAI ("Managed Agents" in error messages and schema descriptions). OpenAI manages the session, orchestration, context compaction, recovery and (optionally) a sandbox; the application supplies tools, input and, optionally, its own execution environment. It is distinct from the Agents SDK (runs in your process) and from the Responses API (single model call). Assistants API (threads/runs) is RETIRED since 2026-08-26 — see
assistants-retired.md.
1. Architecture
flowchart LR
subgraph App["Your application"]
A1[Create agent / session]
A2[Send input events]
A3[Consume SSE stream / webhooks]
A4[Function-tool handler]
A5[Executor codex exec-server<br/>(self_hosted only)]
end
subgraph OpenAI["OpenAI — api.openai.com (OpenAI-Beta: agents=v1)"]
H[Managed Codex harness<br/>model + tool loop + compaction]
subgraph Res["Resources"]
AG[/v1/agents<br/>agent]
SE[/v1/agents/sessions<br/>agent.session]
TU[turns<br/>agent.session.turn]
IT[items<br/>message · reasoning · tool calls]
EV[events<br/>SSE in/out]
SA[subagents<br/>agent.session.subagent]
AR[artifacts<br/>agent.session.artifact]
EN[/v1/agents/environments<br/>agent.environment + files]
ET[/v1/agents/environments/templates<br/>agent.environment.template]
VA[/v1/vaults<br/>vault → vault.credential]
end
SB[(OpenAI-hosted sandbox<br/>/workspace)]
MCP[Remote MCP servers<br/>connection_origin=service]
WS[web_search]
end
A1 --> AG --> SE
A1 --> SE
A2 --> EV --> H
H --> EV --> A3
SE --> TU --> IT
SE --> SA --> TU
SE --> AR
SE -. environment.type=openai_hosted .-> SB
SE -. environment_template_id .-> ET
SE -. vault_ids .-> VA --> MCP
H <--> MCP
H <--> WS
H <--> SB
H -- requires_action: function_call --> A4 -- input.tool_result --> EV
H -- requires_action: environment_connection --> A5 -. wss://codex-cloud-environments.chatgpt.com .-> H
SE --> EN
Pieces (docs/architecture): the harness (OpenAI-hosted Codex instance: model + tool loop + session), the environment (none | openai_hosted | self_hosted), and your application server (submits tasks, receives events, handles function tools, manages self-hosted compute).
Resource tree (spec + live)
agent (POST/GET/POST(update)/DELETE /v1/agents[/{agent_id}]) object "agent"
└─ used by session.agent_id (config copied at creation)
agent.session (POST/GET/POST(update)/DELETE /v1/agents/sessions[/{id}]) object "agent.session"
├─ events POST …/events (input: message | cancel | tool_result → 202) GET …/events?stream=true (SSE, server events)
├─ turns GET …/turns, GET …/turns/{turn_id} object "agent.session.turn"
├─ items GET …/items (14 item types: message, reasoning, function_call(+_output), agent_message, mcp_call, web_search_call, command_execution, *_subagent_call…)
├─ subagents GET …/subagents[/{subagent_id}] (+ /items, /turns[/{turn_id}[/items]]) object "agent.session.subagent"
├─ artifacts GET …/artifacts[/{artifact_id}[/content]], DELETE …/artifacts/{id} object "agent.session.artifact" (openai_hosted, /workspace/outputs)
├─ traces GET …/traces (docs-only, OTLP JSON pages; not in OpenAPI spec; route exists live)
└─ environment (session.environment.id → GET /v1/agents/environments/{id}, GET/POST …/files) object "agent.environment"
agent.environment.template (CRUD /v1/agents/environments/templates[/{id}]) object "agent.environment.template"
vault (POST/GET/DELETE /v1/vaults[/{id}]) object "vault"
└─ vault.credential (POST/GET/POST(rotate)/DELETE /v1/vaults/{id}/credentials[/{cred_id}])Not present anywhere (spec, docs, live): agent versions, budgets objects, memory, deployments, scheduling. The only budget-related surface is the turn error code session_budget_exceeded (see §9).
2. Auth, headers, gating (live-tested)
| Item | Value | Evidence |
|---|---|---|
| Base URL | https://api.openai.com/v1 |
live |
| Auth | Authorization: Bearer $OPENAI_API_KEY |
live |
| Beta header | OpenAI-Beta: agents=v1 required on every /v1/agents/** and /v1/vaults/** call |
Without it: 400 {"type":"invalid_beta","code":"invalid_beta","message":"To access the Agents API, set the 'OpenAI-Beta' header to 'agents=v1'."} |
?beta=true |
not accepted (same 400) | live |
| Restricted-key scopes | api.agents.read, api.agents.write, api.responses.write (inference); api.vaults.read/write for vaults; api.traces.read for trace export |
docs |
| Executor key | separate environment key (dashboard → Agents → Environments → Keys) passed as CODEX_API_KEY; can only connect environments |
docs |
| Streaming | Accept: text/event-stream; POST /sessions with "stream": true or GET …/events?stream=true |
live (Content-Type: text/event-stream) |
| Data | US data residency only, no ZDR (self-hosted sandbox does not make it ZDR-eligible) | docs |
| SDKs | Python/Node client.beta.agents.*, client.beta.agents.sessions.events.stream()/create(), client.beta.agents.vaults.*, client.beta.agents.environments.templates.* (headers added automatically) |
sdk surfaces |
3. Models and pricing
Docs examples use gpt-6-astra ($10 / $50 per 1M in/out, standard). gpt-5.6-terra, gpt-5.6-sol also appear. Live: gpt-5.4-nano rejected — 400 invalid_request_error "The 'gpt-5.4-nano' model is not supported by Managed Agents."; gpt-5.6-luna accepted ($0.20 / $1.20 per 1M) and completed two turns. A no-tool, environment: none turn with instructions "Reply with OK." consumed 6 087 input tokens (harness base instructions) and 6 output tokens ⇒ ≈ $0.0012 per turn on luna; ≈ $0.06 on astra. Billing = model tokens (incl. reasoning as output, prompt-cache rules) + tool rates + container rates for hosted sandboxes. usage on sessions/turns is best-effort and may be null (live: null in list/session, populated on GET …/turns/{id}).
4. Agent definition (POST /v1/agents, CreateAgentParams)
| Field | Type | Notes |
|---|---|---|
model required |
string | requested name preserved |
name |
string ≤128 | nullable |
instructions |
string | appended to the harness's default base instructions (not a replacement) |
reasoning |
{effort: none·minimal·low·medium·high·xhigh·max, summary: concise·detailed·auto} |
omission → model default (live: effort:"medium" returned for luna) |
text |
`{format: {type:text} | {type:json_schema, schema}, verbosity: low·medium·high}` |
service_tier |
auto·default·flex·priority·fast | default auto |
tools[] ≤2000 |
function{name,description,parameters,defer_loading} · tool_search · programmatic_tool_calling{enabled} · `mcp{server_label,transport(http{server_url,headers} |
stdio{command,args,cwd,env_vars}),allowed_tools,required,connection_origin,credential_id,request_metadata}·web_search{mode,context_size,allowed_domains,location}` |
multi_agent |
{enabled, max_concurrent_subagents (default 6)} |
disabled by default |
metadata |
≤16 pairs |
Update semantics (POST /v1/agents/{id}): omitted fields keep values; supplied objects replace the whole field; null resets. Updates affect new sessions only. Response object agent (live example in objects fragment). Delete → {object:"agent.deleted", deleted:true}.
5. Session lifecycle
stateDiagram-v2 [*] --> in_progress: POST /agents/sessions with input (or later input.message) [*] --> idle: POST /agents/sessions without input (self_hosted / openai_hosted) idle --> in_progress: agent.session.input.message in_progress --> requires_action: function_call / environment_connection requires_action --> in_progress: input.tool_result / executor connects in_progress --> idle: turn completed | failed | cancelled in_progress --> failed: agent.session.failed idle --> [*]: DELETE (object agent.session.deleted; 409 while busy → retry)
- Create (
CreateAgentSessionParams):agent(inlineSessionAgentConfigParam) and/oragent_id(override fields merge for that session only; arrays replace),environmentrequired (none|openai_hosted{…}|self_hosted{workspace_directory, capability_directories}),vault_ids[],input(string shorthand or[{role:user, content:[input_text|input_image]}], required whenenvironment.type: none),stream(SSE for the first turn),metadata. Success201(live) — spec says 201. - Turn = one cycle of work; a message to an idle session starts a turn, a message during an active turn steers it. Turn statuses:
queued → in_progress → waiting → completed | failed | cancelled. Cancel withagent.session.input.cancel(live: 202 even when idle). - Update (
POST …/sessions/{id}): onlyagent.model,agent.reasoning.effort,agent.service_tier,metadata(live: metadata-only update 200). Cannot changeinstructions,tools,text,reasoning.summary,multi_agent— create a new session. - Required actions (
session.required_actions[]):function_call{turn_id,call_id,name,arguments}→ answer withagent.session.input.tool_result{turn_id,call_id,success,output|error};environment_connection{environment_id}→ start/reconnect the executor. API waits up to 5 minutes for an executor. - Items (saved history,
GET …/items?order=asc):message(role user/assistant,phase: commentary|final_answer),reasoning,function_call,function_call_output,mcp_call,web_search_call,command_execution{command,cwd,exit_code,output,duration_ms},agent_message(inter-agent),create_subagent_call,send_subagent_input_call,resume_subagent_call,wait_for_subagents_call,interrupt_subagent_call,close_subagent_call. Live: 2 items after one turn (userinput_text+ assistantoutput_text"OK.",phase:"final_answer"). - Pagination everywhere:
limit(1–100, default 20),order(asc|desc),after= previouslast_id; list bodies{object:"list", data, first_id, last_id, has_more}(environment files use{object:"page", data, next, has_more}).
6. Streaming events (SSE, server → client)
POST /sessions (stream:true) or GET /sessions/{id}/events?stream=true. Each SSE event: line equals the JSON type; every payload has event_id. No replay after disconnect — recover by re-opening the stream, retrieving session + items, then applying buffered updates by item_id. Live sequence for a no-tool turn (identical for the first turn via POST and a follow-up via GET):
agent.session.created (POST only) → agent.session.turn.created → agent.session.turn.item.added (user message)
→ agent.session.in_progress → agent.session.turn.in_progress → agent.session.turn.item.added (assistant message)
→ agent.session.turn.content_part.added → agent.session.turn.output_text.delta ×N → agent.session.turn.output_text.done
→ agent.session.turn.content_part.done → agent.session.turn.item.done → agent.session.turn.completedAll 31 spec event types (grouped): session created · in_progress · idle · requires_action · failed; turn created · in_progress · completed · failed · cancelled (carry turn + usage); item turn.item.added · turn.item.done; text turn.content_part.added/done · turn.output_text.delta/done; reasoning turn.reasoning_summary_part.added/done · turn.reasoning_summary_text.delta/done; command agent.output.command_execution_output.delta; environment environment.pending · ready · connected · disconnected · failed · reset; subagent subagent.created · active · closed; error. Terminal for a turn: turn.completed | turn.failed | turn.cancelled — agent.session.idle alone is not success; a completed turn may still contain failed tool calls. Subagent turn events (turn.subagent_id != null) do not end the root stream. Live observation: agent.session.idle follows turn.completed (seen when the SSE connection was kept open — examples/openai/agents/send_turn_stream.sh); the stream does not close by itself after a turn, so clients must stop on the terminal turn event.
Observed usage (gpt-5.6-luna, environment: none, "Reply with OK."): turn 1 input_tokens 6087 / cached 0 / output 5, turn 2 6102 / cached 6084 / 5, turn 3 6117 / cached 6099 / 5 — the ~6k-token harness prompt is prompt-cached from the second turn. usage was null on list/turn-event payloads and populated on GET …/turns/{id} shortly after completion (best-effort accounting).
Client → server input events (POST …/events, 202 empty body, optional Idempotency-Key): agent.session.input.message{input[]}, agent.session.input.cancel, agent.session.input.tool_result.
Webhooks (signed POST, verify with SDK webhooks.verify_signature): agent.session.created (self-hosted payload includes environment_id, environment_type, connect.remote_url), agent.session.action_required (required_action.type only — retrieve the session for details), agent.session.in_progress, agent.session.idle, agent.session.failed. No deletion webhook.
7. Environments (summary; details in agents-environments-and-vaults.md)
environment.type |
Who runs compute | Files/artifacts | Tools available |
|---|---|---|---|
none |
nobody — model + remote MCP + function tools only; initial input required |
none (read items) | no Bash/apply_patch, no executor MCP |
openai_hosted |
OpenAI sandbox (Linux, Python, Node; /workspace) |
live files API; /workspace/outputs published as artifacts after each turn |
everything incl. stdio MCP (network enabled required) |
self_hosted |
you (codex exec-server --remote <remote_url> --environment-id <id> with CODEX_API_KEY=environment key) |
your provider's storage; no artifacts | everything; private network |
8. Credentials, subagents, artifacts
- Vaults hold MCP credentials (
static_bearer,mcp_oauthwith optional refresh,environment_variablefor hosted sandboxes via placeholder + proxy). Attach withvault_ids; credential selected bymcp_server_urlmatch orcredential_id. Secrets are write-only. Vaults apply only toconnection_origin: service. - Subagents: enable
agent.multi_agent.enabled; the harness adds create/send/wait/interrupt/close tools (not declared by you). Subagents inherit MCP tools + web search, share the environment, cannot use function tools. Observe viaagent.session.subagent.*events and*_subagent_callitems; per-subagent items/turns endpoints.turn.subagent_idisnullfor the root agent. - Artifacts (
agent.session.artifact): immutable copies of files under/workspace/outputspublished when a hosted turn completes;GET …/contentstreams bytes (application/octet-stream); one file per request (ask the agent to ZIP). Limits: 50 files/creation request, inline 5 MiB/file & 10 MiB total, Files-API copy 50 MiB, artifact 200 MiB, outputs 500 MiB/turn. Deleting an artifact leaves the environment file intact.
9. Errors and limits
- Envelope:
{"error":{"type","code","message","param"}}. Live:invalid_beta(400),invalid_request_error(400, unsupported model),not_found_error(404, "No managed agent session found: …", "No managed agent resource found: …", "No vault resource found: …"). Spec lists 400/401/403/404/409/500/503 for most operations (ErrorResponse-2). - Turn error codes (
turn.error.code):context_length_exceeded,session_budget_exceeded,usage_limit_exceeded,credit_balance_exhausted,rate_limit_exceeded,server_overloaded,cyber_policy,connection_failed,server_error,authentication_error,invalid_request,resource_not_found,sandbox_error,executor_version_incompatible,active_turn_not_steerable,request_timeout,internal_error. - Sandbox expiry: hosted sandbox deletable after ~1 h without activity/keep-alives (not configurable). Delete may return
409while setup/execution finishes. - No account-specific rate limits documented; response headers observed:
x-request-id,openai-processing-ms,openai-version: 2020-10-01(nox-ratelimit-*on Agents calls).
10. Relation to Responses API, Agents SDK, ChatKit
| Agents API | Agents SDK | Responses API | |
|---|---|---|---|
| Where the loop runs | OpenAI (managed Codex harness) | your process | your code (single calls) |
| State | server-side session, turns, items | your storage / SDK sessions | manual chaining or Conversations |
| Tools | function (you run), MCP (OpenAI or environment), web_search, programmatic tool calling, tool_search, sandbox shell | anything in your app | hosted + your tools |
| Inference billing | api.responses.write scope — model calls billed at Responses rates |
same models | same |
Function-tool implementations from Responses are reusable; content types (input_text, output_text, input_image) and reasoning/text/service_tier settings mirror Responses. ChatKit is an embeddable UI (its API is sessions/threads for Agent Builder workflows — see chatkit.md). Agent Builder shuts down 2026-11-30.
11. Live verification log (2026-09-18, total ≈ $0.003)
| Call | Result |
|---|---|
GET /v1/agents (no header / ?beta=true) |
400 invalid_beta |
GET /v1/agents, /agents/sessions, /agents/environments/templates, /vaults, /chatkit/threads |
200 empty lists |
POST /v1/agents gpt-5.4-nano |
400 "not supported by Managed Agents" |
POST /v1/agents gpt-5.6-luna |
201 agent |
POST /v1/agents/sessions (agent_id, env none, "Reply with OK.", stream) |
201 SSE, 13 events, "OK." |
GET session / items / turns / turn / subagents / artifacts |
200 |
GET …/events?stream=true + POST …/events (message, Idempotency-Key) |
200 stream (12 events) / 202 |
POST …/sessions/{id} metadata · POST …/events cancel (idle) |
200 · 202 |
DELETE session · agent · GET deleted session |
200 · 200 · 404 |
GET …/traces, /agents/environments/{bogus}, /vaults/{bogus} |
404 not_found_error (routes exist) |
POST/GET/POST/DELETE /v1/agents/environments/templates[/{id}] (examples) |
201 / 200 / 200 / 200 — config only, free |
GET /v1/agents/{agent_id} (examples) |
200 |
18 example files (examples/openai/agents/*.{sh,py,ts}) |
all ran; 6 more cheap turns, prompt-cached from turn 2 |
GET /v1/assistants, /v1/threads/{x} |
404 empty body (retired) |
Raw sanitized captures: tmp-live/agents/*.json; request log: reports/live-requests.jsonl.