SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
5.7 KB

# OpenAI multimodal input (images, files/PDFs, audio) — Responses and Chat Completions

Status: DOCUMENTED + LIVE_VERIFIED for image input (1×1 PNG data URL on gpt-4.1-nano, both APIs) and for token counting of an image part; file and audio input DOCUMENTED only. Machine-readable: input[](message).content[](input_image|input_file|input_text) rows in generated/fragments/parameters/openai-responses.json; messages[](user).content[](image_url|file|input_audio|text) rows in openai-chat-completions.json; content-part objects in generated/fragments/objects/openai-responses-objects.json.

Sources

Last verified: 2026-09-18

# 1. Content-part shapes

Modality Responses input[].content[] Chat messages[].content[]
Text {"type":"input_text","text":"…"} {"type":"text","text":"…"}
Image {"type":"input_image","image_url":"https://… | data:image/png;base64,…","detail":"low|high|auto|original"} or {"type":"input_image","file_id":"file_…","detail":…} {"type":"image_url","image_url":{"url":"https://… | data:…","detail":"auto|low|high"}}
File / PDF {"type":"input_file","file_id":"file_…"} · {"type":"input_file","file_url":"https://…pdf"} · {"type":"input_file","filename":"doc.pdf","file_data":"data:application/pdf;base64,…"}, optional "detail":"auto|low|high" (PDF page images) {"type":"file","file":{"file_id":"…"}} or {"type":"file","file":{"filename":"doc.pdf","file_data":"data:…"}}
Audio {"type":"input_audio","input_audio":{"data":"<base64>","format":"mp3|wav"}} (audio-capable models) {"type":"input_audio","input_audio":{"data":"<base64>","format":"wav|mp3"}} + modalities/audio for audio output (gpt-audio-1.5, gpt-4o-audio-preview)
Cache breakpoint any part may carry "prompt_cache_breakpoint":{"mode":"explicit"} (GPT-5.6+) same on text parts

include:["message.input_image.image_url"] returns image URLs back in stored items. Responses detail is required in the spec schema for input_image but defaults to auto when omitted in practice (docs).

# 2. Image requirements and token cost

  • Formats: PNG, JPEG, WEBP, non-animated GIF. Up to 512 MB per request, 1,500 images per request; patch-based models reject images >30,000 patches after resizing (not auto-resized).
  • detail: low (coarse; not always cheaper on patch models), high, original (dense/OCR/computer-use, supported on gpt-6-astra and recent models), auto (model default).
  • Patch-based models (gpt-5.x, gpt-4.1-mini/nano, o4-mini…): tokens = ceil(w/32)·ceil(h/32) after fitting the detail's pixel limit and patch budget, × multiplier (1.2 for gpt-5.x/gpt-6-astra, 1.62 gpt-4.1-mini, 2.46 gpt-4.1-nano, 1.5 gpt-5-nano, 1.72 o4-mini). Example gpt-6-astra high: 1024×1024 → 1,229 tokens; 2048×2048 → 3,000; 4096×512 → 2,458.
  • Tile-based models (gpt-4o, gpt-4.1, gpt-5, gpt-5.1, o1/o3): base + tiles (gpt-4o/4.1: 85 + 170 per 512 px tile after fitting 2048² and 768 px short side; gpt-4o-mini 2833 + 5667; gpt-5/5.1: 70 + 140). detail:low = base tokens only.
  • Use POST /v1/responses/input_tokens to get the exact count before sending (live: 1×1 PNG at detail:low + text + tool = 49 tokens).
  • Limitations: medical images, non-Latin text, rotated text, small text (use original), spatial precision, counting, panoramas, CAPTCHAs blocked, metadata/filenames ignored.

# 3. Files / PDFs

  • Sources: file_id (upload with purpose:"user_data"), file_url (PDF only), file_data base64 data URL (+filename). Non-PDF office/text files supported for file_id/file_data (docx, pptx, xlsx, csv, txt, md, json, html, rtf…; see accepted-types table in the guide), but image/chart extraction only for PDFs.
  • Limits: each file < 50 MB and combined 50 MB per request; PDF parsing puts extracted text and page images in context (detail controls page-image fidelity) → higher token usage. Requires vision-capable models (gpt-4o and later).
  • Token counting endpoint supports files; the Responses spec calls these "file inputs" (guide URL /guides/file-inputs).

# 4. Audio

  • Chat Completions: modalities:["text","audio"], audio:{voice, format} for output; input via input_audio parts; assistant turns reference prior audio with audio:{id}; usage adds audio_tokens in both details objects. Models: gpt-audio-1.5, gpt-4o-audio-preview, gpt-4o-mini-audio-preview.
  • Responses: input_audio part exists in the schema and the SSE union has response.audio.delta/done, response.audio.transcript.delta/done, but the Responses docs currently describe text/image inputs only — treat audio-in on Responses as UNVERIFIED; use Chat Completions for the audio-chat pattern or the Realtime API.

# 5. Live results (2026-09-18)

Call Result
Responses gpt-4.1-nano, input_text + input_image 1×1 PNG data URL detail:low, 32 tokens 200 completed, input_tokens 15, output_tokens 2
Chat gpt-4.1-nano, text + image_url data URL detail:low 200, prompt_tokens 13, completion_tokens 1
POST /v1/responses/input_tokens with instructions + image + function tool (gpt-5.4-nano) input_tokens 49

Both requests cost < $0.000003. A 1×1 image costs ≈ 3–5 tokens on gpt-4.1-nano (1 patch × 2.46 → 3, plus part framing).