OpenAI multimodal input (images, files/PDFs, audio) — Responses and Chat Completions
Status: DOCUMENTED + LIVE_VERIFIED for image input (1×1 PNG data URL on gpt-4.1-nano, both APIs) and for token counting of an image part; file and audio input DOCUMENTED only. Machine-readable: input[](message).content[](input_image|input_file|input_text) rows in generated/fragments/parameters/openai-responses.json; messages[](user).content[](image_url|file|input_audio|text) rows in openai-chat-completions.json; content-part objects in generated/fragments/objects/openai-responses-objects.json.
Sources
- https://developers.openai.com/api/docs/guides/images-vision (requirements, detail levels, sizing, token math) · https://developers.openai.com/api/docs/guides/file-inputs · https://developers.openai.com/api/docs/guides/audio-chat-completions · https://developers.openai.com/api/docs/guides/token-counting
- OpenAPI
InputContent(InputTextContent,InputImageContent,InputFileContent),InputAudio,ChatCompletionRequestMessageContentPart{Text,Image,Audio,File}
Last verified: 2026-09-18
1. Content-part shapes
| Modality | Responses input[].content[] |
Chat messages[].content[] |
|---|---|---|
| Text | {"type":"input_text","text":"…"} |
{"type":"text","text":"…"} |
| Image | {"type":"input_image","image_url":"https://… | data:image/png;base64,…","detail":"low|high|auto|original"} or {"type":"input_image","file_id":"file_…","detail":…} |
{"type":"image_url","image_url":{"url":"https://… | data:…","detail":"auto|low|high"}} |
| File / PDF | {"type":"input_file","file_id":"file_…"} · {"type":"input_file","file_url":"https://…pdf"} · {"type":"input_file","filename":"doc.pdf","file_data":"data:application/pdf;base64,…"}, optional "detail":"auto|low|high" (PDF page images) |
{"type":"file","file":{"file_id":"…"}} or {"type":"file","file":{"filename":"doc.pdf","file_data":"data:…"}} |
| Audio | {"type":"input_audio","input_audio":{"data":"<base64>","format":"mp3|wav"}} (audio-capable models) |
{"type":"input_audio","input_audio":{"data":"<base64>","format":"wav|mp3"}} + modalities/audio for audio output (gpt-audio-1.5, gpt-4o-audio-preview) |
| Cache breakpoint | any part may carry "prompt_cache_breakpoint":{"mode":"explicit"} (GPT-5.6+) |
same on text parts |
include:["message.input_image.image_url"] returns image URLs back in stored items. Responses detail is required in the spec schema for input_image but defaults to auto when omitted in practice (docs).
2. Image requirements and token cost
- Formats: PNG, JPEG, WEBP, non-animated GIF. Up to 512 MB per request, 1,500 images per request; patch-based models reject images >30,000 patches after resizing (not auto-resized).
detail:low(coarse; not always cheaper on patch models),high,original(dense/OCR/computer-use, supported on gpt-6-astra and recent models),auto(model default).- Patch-based models (gpt-5.x, gpt-4.1-mini/nano, o4-mini…): tokens = ceil(w/32)·ceil(h/32) after fitting the detail's pixel limit and patch budget, × multiplier (1.2 for gpt-5.x/gpt-6-astra, 1.62 gpt-4.1-mini, 2.46 gpt-4.1-nano, 1.5 gpt-5-nano, 1.72 o4-mini). Example gpt-6-astra high: 1024×1024 → 1,229 tokens; 2048×2048 → 3,000; 4096×512 → 2,458.
- Tile-based models (gpt-4o, gpt-4.1, gpt-5, gpt-5.1, o1/o3): base + tiles (gpt-4o/4.1: 85 + 170 per 512 px tile after fitting 2048² and 768 px short side; gpt-4o-mini 2833 + 5667; gpt-5/5.1: 70 + 140).
detail:low= base tokens only. - Use
POST /v1/responses/input_tokensto get the exact count before sending (live: 1×1 PNG atdetail:low+ text + tool = 49 tokens). - Limitations: medical images, non-Latin text, rotated text, small text (use
original), spatial precision, counting, panoramas, CAPTCHAs blocked, metadata/filenames ignored.
3. Files / PDFs
- Sources:
file_id(upload withpurpose:"user_data"),file_url(PDF only),file_database64 data URL (+filename). Non-PDF office/text files supported forfile_id/file_data(docx, pptx, xlsx, csv, txt, md, json, html, rtf…; see accepted-types table in the guide), but image/chart extraction only for PDFs. - Limits: each file < 50 MB and combined 50 MB per request; PDF parsing puts extracted text and page images in context (
detailcontrols page-image fidelity) → higher token usage. Requires vision-capable models (gpt-4o and later). - Token counting endpoint supports files; the Responses spec calls these "file inputs" (guide URL
/guides/file-inputs).
4. Audio
- Chat Completions:
modalities:["text","audio"],audio:{voice, format}for output; input viainput_audioparts; assistant turns reference prior audio withaudio:{id}; usage addsaudio_tokensin both details objects. Models:gpt-audio-1.5,gpt-4o-audio-preview,gpt-4o-mini-audio-preview. - Responses:
input_audiopart exists in the schema and the SSE union hasresponse.audio.delta/done,response.audio.transcript.delta/done, but the Responses docs currently describe text/image inputs only — treat audio-in on Responses asUNVERIFIED; use Chat Completions for the audio-chat pattern or the Realtime API.
5. Live results (2026-09-18)
| Call | Result |
|---|---|
Responses gpt-4.1-nano, input_text + input_image 1×1 PNG data URL detail:low, 32 tokens |
200 completed, input_tokens 15, output_tokens 2 |
Chat gpt-4.1-nano, text + image_url data URL detail:low |
200, prompt_tokens 13, completion_tokens 1 |
POST /v1/responses/input_tokens with instructions + image + function tool (gpt-5.4-nano) |
input_tokens 49 |
Both requests cost < $0.000003. A 1×1 image costs ≈ 3–5 tokens on gpt-4.1-nano (1 patch × 2.46 → 3, plus part framing).