SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
10.3 KB

# OpenAI Images API (/v1/images/*)

Status: POST /v1/images/generations DOCUMENTED + LIVE_VERIFIED · POST /v1/images/edits DOCUMENTED (UNVERIFIED, paid) · POST /v1/images/variations LEGACY + RETIRED (live 404) · DALL·E 2/3 RETIRED (removed 2026-05-12) · gpt-image-1 / 1-mini / 1.5 / chatgpt-image-latest DEPRECATED (shutdown 2026-10-23 / 2026-12-01). Sources: Images reference · Generation streaming events · Edit streaming events · Image generation guide · Pricing · Deprecations · model pages gpt-image-*, chatgpt-image-latest, dall-e-3 · OpenAPI CreateImageRequest, CreateImageEditRequest, CreateImageVariationRequest, ImageGenStreamEvent, ImageEditStreamEvent. Last verified: 2026-09-18 (live calls logged in reports/live-requests.jsonl, raw in tmp-live/media-agent-batch2.json). Machine-readable: generated/fragments/endpoints/openai-images-video-embeddings-moderation.json, parameters/openai-images.json, streaming-events/openai-images.json, objects/openai-media-objects.json, prices/openai-media.json.

The Image API is the direct way to generate/edit images. The image_generation tool of the Responses API (multi-turn editing, file IDs, mainline model + tool.model) is documented by the tools agent — see docs/tools/.

# Endpoints

Method / path Purpose Body Streaming Status
POST /v1/images/generations text → image(s) JSON yes (stream, partial_images) LIVE_VERIFIED
POST /v1/images/edits 1–16 source images (+ optional mask) + prompt → image(s) multipart/form-data (image, image[], mask) or JSON (images[] with file_id/image_url, mask{}) yes DOCUMENTED, UNVERIFIED
POST /v1/images/variations variations of a square PNG, dall-e-2 only multipart no RETIRED — live 2026-09-18: HTTP 404, empty body

SDKs: Python client.images.generate / edit / create_variation; Node client.images.generate / edit / createVariation. Batch API: gpt-image-2, 1.5, 1, 1-mini, chatgpt-image-latest (not 2.5 per model pages).

# Model matrix

Model Snapshot / alias Gen Edit Sizes Quality Background Formats input_fidelity Status
gpt-image-2.5-sunburst (editing precision) -2026-09-08 ✓ ✓ any WxH (rules below) + auto low/medium/high/xhigh/max/auto transparent/opaque/auto png/jpeg/webp high/low DOCUMENTED (released 2026-09-08)
gpt-image-2.5-flare (fast) -2026-09-08 ✓ ✓ any WxH + auto low/medium/high/xhigh/max/auto transparent/opaque/auto png/jpeg/webp high/low DOCUMENTED
gpt-image-2 -2026-04-21 ✓ ✓ any WxH + auto low/medium/high/auto transparent PREVIEW/opaque/auto png/jpeg/webp always high (param not settable) DOCUMENTED
gpt-image-1.5 -2025-12-16 ✓ ✓ 1024², 1536×1024, 1024×1536, auto low/medium/high/auto ✓ png/jpeg/webp high/low DEPRECATED → 2026-12-01
chatgpt-image-latest — ✓ ✓ same as 1.5 low/medium/high/auto ✓ png/jpeg/webp high/low DEPRECATED → 2026-12-01
gpt-image-1 — ✓ ✓ same low/medium/high/auto ✓ png/jpeg/webp high/low DEPRECATED → 2026-10-23
gpt-image-1-mini — ✓ ✓ same low/medium/high/auto ✓ png/jpeg/webp not supported DEPRECATED → 2026-12-01 · LIVE_VERIFIED
dall-e-3 — ✓ ✗ 1024², 1792×1024, 1024×1792 standard/hd (+style vivid/natural, n=1) ✗ url / b64_json ✗ RETIRED (live: The model 'dall-e-3' does not exist.)
dall-e-2 — ✓ ✓ (1 PNG < 4 MB) 256², 512², 1024² standard ✗ url / b64_json ✗ RETIRED

Live /v1/models (2026-09-18) lists all GPT image ids above (incl. dated snapshots) but no dall-e-*.

Arbitrary sizes (gpt-image-2 / 2.5): WIDTHxHEIGHT, both multiples of 16, aspect ratio between 1:3 and 3:1, max edge 3840 px, total pixels 655,360–8,294,400 (4K); > 2560×1440 is experimental. Popular: 1024², 1536×1024, 1024×1536, 2048², 2048×1152, 3840×2160, 2160×3840. Older GPT image models reject anything else — live: size=123x456 on gpt-image-1-mini → 400 {type: image_generation_user_error, param: size, code: invalid_value, message: "Invalid size '123x456'. Supported sizes are 1024x1024, 1024x1536, 1536x1024, and auto."}.

# Parameters (generations)

Param Type Default Models Notes
prompt * string — all ≤ 32 000 chars GPT image; 1 000 dall-e-2; 4 000 dall-e-3
model string dall-e-2 (docs) all DALL·E removed → always pass a model. Enum in spec lacks chatgpt-image-latest for generations (model page says supported).
n int 1–10 1 all dall-e-3: 1 only
quality enum auto see matrix standard/hd are DALL·E values
size string auto see matrix
background transparent/opaque/auto auto GPT image use png/webp with transparent
output_format png/jpeg/webp png GPT image jpeg is fastest
output_compression 0–100 100 GPT image (jpeg/webp)
moderation low/auto auto GPT image low = less restrictive filtering
stream bool false GPT image SSE events below
partial_images 0–3 0 GPT image +100 output tokens per partial
response_format url/b64_json url DALL·E only live on gpt-image-1-mini → 400 unknown_parameter
style vivid/natural vivid dall-e-3 only live on gpt-image-1-mini → 400 unknown_parameter
user string — all safety identifier (see docs/openai/safety.md)

# Edits: inputs, masks, fidelity

  • Inputs: multipart image (single) or repeated image[] (≤ 16 images, png/webp/jpg, each < 50 MB for GPT image); JSON body alternative images: [{file_id | image_url}] (URL or base64 data URL). dall-e-2: one square PNG < 4 MB.
  • Mask: PNG with alpha channel; fully transparent pixels = area to edit; same size/format as the first image; applied to the first image only. With GPT Image, masking is prompt-guided (not pixel-exact). JSON form: mask: {file_id | image_url}.
  • input_fidelity low (default) | high: preserve details/faces of inputs. gpt-image-1, 1.5, chatgpt-image-latest, 2.5 ✓; gpt-image-1-mini ✗; gpt-image-2 always high (omit the param; more input tokens on edits). Changelog notes a fixed bug where 1.5/chatgpt-image-latest ignored low.
  • Spec default model for edits = gpt-image-1.5. Same quality/size/background/output_*/moderation/stream/partial_images/n/user semantics as generations.

# Response object (ImagesResponse)

json
{"created": 1789782224, "background": "opaque", "output_format": "jpeg", "quality": "low", "size": "1024x1024",
 "data": [{"b64_json": "<base64 jpeg>", "generation_id": "<opaque id>"}],
 "usage": {"input_tokens": 10, "input_tokens_details": {"image_tokens": 0, "text_tokens": 10},
           "output_tokens": 272, "output_tokens_details": {"image_tokens": 272, "text_tokens": 0}, "total_tokens": 282}}

Live (gpt-image-1-mini, low, 1024², jpeg, 5.7 s, 39 383 bytes). generation_id in each data[] item is undocumented (LIVE_DISCOVERED). url/revised_prompt were DALL·E-only. usage is present for all GPT image models although the reference text still says "gpt-image-1 only".

# Streaming (GPT image models only)

stream: true → SSE stream of event:/data: pairs, then connection close. Events (generated/fragments/streaming-events/openai-images.json):

Event When Payload
image_generation.partial_image each partial (≤ partial_images, may be fewer) b64_json, partial_image_index, background, created_at, output_format, quality, size, type
image_generation.completed final same fields minus index + usage
image_edit.partial_image / image_edit.completed same for /edits same shapes

Streaming not exercised live (UNVERIFIED; cost).

# Pricing pointers (USD, standard tier; batch = 50 %)

Model Text in / cached Image in / cached Image out Text out
gpt-image-2.5-sunburst / -flare / gpt-image-2 $5 / $1.25 per 1M $8 / $2 $30 —
gpt-image-1.5 / chatgpt-image-latest $5 / $1.25 $8 / $2 $32 $10
gpt-image-1 $5 / $1.25 $10 / $2.5 $40 —
gpt-image-1-mini $2 / $0.2 $2.5 / $0.25 $8 —

Legacy output-token counts (pre-gpt-image-2): low 272 / 408 / 400 tokens, medium 1056 / 1584 / 1568, high 4160 / 6240 / 6208 for 1024² / 1024×1536 / 1536×1024 → per-image tables on the model pages (e.g. gpt-image-1-mini low 1024² ≈ $0.005; 1.5 high portrait $0.20; gpt-image-2 low 1024² $0.006, high $0.211). Our live call: 272 output tokens × $8/M + 10 text tokens ≈ $0.0022. gpt-image-2/2.5 token counts vary by size/quality (use usage; the guide's calculator does not cover 2.5). Each partial image adds 100 output tokens. Rate limits (all GPT image, model pages): Tier 1 100k TPM / 5 IPM … Tier 5 8M TPM / 250 IPM.

# Limits and errors

  • Latency up to ~2 min for complex prompts; text rendering, consistency and layout precision remain weak spots (guide).
  • All prompts/images are filtered per usage policies; moderation: "low" relaxes filtering. Org verification may be required for GPT image models.
  • User-correctable failures: error.type = "image_generation_user_error" with stable error.code (e.g. invalid_value) — do not retry unchanged. Unknown params → invalid_request_error / unknown_parameter.
  • Data residency: /v1/images/generations and /edits are available on all regional endpoints (gpt-image-2/2.5/1/1.5/1-mini).

# Examples and tests

examples/openai/images/ (generate.sh|py|ts, edit.sh UNVERIFIED, stream.py UNVERIFIED) · tests/openai/test_images.py (marker run_image_tests; the validation-error test is free and unmarked).