Realtime, voice and media — what each of the four providers exposes
Status: endpoints and statuses from generated/endpoints.json (OpenAI realtime 14, live 10, audio 10, images 3, videos 10, embeddings 1, moderations 1; xAI voice 17, images 2, videos 4, embeddings 1; Gemini live 3, music-generation 1, video-generation 1, image-generation 1 (Imagen, RETIRED), embeddings 2, transcription 1, auth-tokens 1), models from generated/models.json, prices from generated/pricing.json, message families from generated/streaming-events.json; Anthropic side from generated/models.json modalities (text, image, pdf → text for every Claude model). Live: OpenAI client secrets / WebSocket text turn / TTS / STT / image generation / embeddings / moderation LIVE_VERIFIED; xAI realtime text turn, client secrets, TTS, STT, image generation + edit, video generation LIVE_VERIFIED; Gemini Live WebSocket (setup + text + audio), ephemeral tokens, TTS, transcription, embeddings, Lyria RealTime LIVE_VERIFIED, image / Veo / Lyria 3.x generation ACCOUNT_RESTRICTED on this free-tier key.
Sources: https://developers.openai.com/api/docs/guides/realtime · …/guides/audio · …/guides/image-generation · …/guides/embeddings · https://platform.claude.com/docs/en/build-with-claude/vision · https://docs.x.ai/developers/model-capabilities/audio/speech-to-speech · https://docs.x.ai/developers/model-capabilities/image/generation · https://docs.x.ai/developers/model-capabilities/video/generation · https://ai.google.dev/gemini-api/docs/live · https://ai.google.dev/gemini-api/docs/speech-generation · https://ai.google.dev/gemini-api/docs/image-generation · https://ai.google.dev/gemini-api/docs/video · https://ai.google.dev/gemini-api/docs/music-generation · https://ai.google.dev/gemini-api/docs/embeddings · docs/openai/{realtime,live,audio,images,video,embeddings,moderation}.md · docs/xai/{voice,images,videos}.md · docs/gemini/{live-api,live-events,speech-generation,transcription,image-generation,video-generation,music-generation,embeddings,multimodal-input}.md
Last verified: 2026-09-18
1. Modality coverage
| Modality | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Text in / out | all text models | all Claude models | all Grok text models | all Gemini / Gemma text models |
| Image in | all GPT-5.x / 6 / o-series / 4.x (input_image), image models, realtime 2.x |
all Claude models (image block; 100 / 600 per request) |
all Grok text models (input_image / image_url, ≥512 px; billed at the text input price) + Imagine edit inputs |
all Gemini 3.x / 2.5 text models (inlineData / fileData; mediaResolution levels), image models, Live, embedding-2 |
| PDF / documents in | input_file |
document block with citations |
input_file (attachment search $10/1k) |
inlineData {application/pdf} / fileData (DOCUMENT modality tokens); URL context for remote PDFs |
| Audio in | Chat audio models, Realtime, Live, transcription | — | realtime session, STT (POST /v1/stt, WS), grok-imagine-video-1.5 reference audio |
native on every 3.x text model (≈25–32 tok/s), Live, transcribe models, embedding-2 |
| Video in | — | — | grok-imagine-video edits/extensions; view_x_video sub-tool |
native on 3.x text models (frames + audio, videoMetadata {fps}, YouTube URLs), Omni, embedding-2 |
| Audio out | POST /v1/audio/speech, Chat modalities, Realtime, Live |
— | POST /v1/tts, wss://api.x.ai/v1/tts, realtime session |
TTS models (responseModalities: [AUDIO]), Live, live-translate |
| Image out | POST /v1/images/generations|edits, image_generation tool |
— | POST /v1/images/generations|edits, image_generation tool |
image models via generateContent (responseModalities: [TEXT, IMAGE]), Interactions response_format {type: image} |
| Video out | POST /v1/videos (Sora 2) — DEPRECATED, shutdown 2026-09-24 |
— | POST /v1/videos/generations|edits|extensions (Imagine video) |
Veo 3.1 (:predictLongRunning), gemini-omni-1.1-flash (Interactions) |
| Music out | — | — | — | Lyria 3.5 / 3 previews (generateContent), Lyria RealTime WebSocket |
| Embeddings | POST /v1/embeddings |
— | POST /v1/embeddings — documented, 404 for this team (grok-embedding-small, unpriced) |
:embedContent / :batchEmbedContents (gemini-embedding-2, multimodal) |
| Moderation | POST /v1/moderations, inline moderation |
— (built-in refusals) | — (usage-guideline violations billed; respect_moderation flags) |
— (safetySettings thresholds, safetyRatings) |
| Provenance | POST /v1/content_provenance_checks |
C2PA credentials on generated media | — | SynthID on all generated media (C2PA on Nano Banana 2 Lite); no check endpoint |
Anthropic's answer to every generation row is "none": Claude produces text only; media generation is delegated to a tool you run. Gemini is the only provider whose text models natively take audio and video.
2. Realtime voice
| Aspect | OpenAI Realtime API (GA) | OpenAI Live API (GPT-Live) | xAI Speech-to-Speech | Gemini Live API |
|---|---|---|---|---|
| Models | gpt-realtime-2.1 ($32/$64 audio, $4/$24 text, $5 image per 1M), -2.1-mini ($10/$20), gpt-realtime-2, -1.5, gpt-realtime(-mini) (→ 2027-01-20), gpt-realtime-translate, gpt-realtime-whisper; 128k ctx |
gpt-live-1 ($0.05 / min, billed per second), gpt-live-transcribe ($0.017 / min) |
grok-voice-think-fast-2.0 = grok-voice-latest (since 2026-08-05; -1.0 LEGACY); $0.08 / min ($4.80 / h) of audio sent or received + $0.004 per text conversation.item.create; built-in "think-fast" reasoning (reasoning.effort: high|none); not in GET /v1/models |
gemini-3.8-live (native audio, interleaved thinking, GA 2026-09-15), gemini-3.8-live-extended-thinking (LOW/MEDIUM/HIGH background reasoning, speaks fillers), gemini-3.1-flash-live-preview (legacy preview), gemini-2.5-flash-native-audio-preview-12-2025; text $0.75 / $4.50, audio $3 in (≈ $0.005 / min) / $12 out (≈ $0.018 / min), image-video $1 per 1M; free tier; gemini-3.5-live-translate-preview (speech-to-speech translation, 70+ languages), gemini-3.5-transcribe-live (streaming STT, 10-min sessions) |
| Transports | WebSocket wss://api.openai.com/v1/realtime?model=, WebRTC POST /v1/realtime/calls, SIP sip:<PROJECT_ID>@sip.api.openai.com (+ accept/reject/refer/hangup) |
WebRTC POST /v1/live/sessions, WebSocket WS /v1/live/sessions, fork/attach/SIP, recordings |
WebSocket wss://api.x.ai/v1/realtime?model=grok-voice-latest[&reasoning.effort=…][&conversation_id=][&call_id=]; SIP via POST /v2/phone-numbers {origin: byo_trunk, phone_number, webhook, sip_auth} → sip.voice.x.ai, POST /v1/realtime/calls/{id}/refer|hangup, webhook realtime.call.incoming; no WebRTC |
WebSocket wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (API key) or …BidiGenerateContentConstrained?access_token= (ephemeral token); no WebRTC/SIP |
| Session config | session.update {session:{type, model, instructions, audio{input{format, transcription, noise_reduction, turn_detection}, output{voice, speed}}, tools, tool_choice, max_output_tokens, tracing}} |
session {model, instructions, audio, delegation: client | responses{model, tools…}, store, input[]} |
session.update {instructions, voice (28 built-in, default eve; server reported xai_ara), reasoning.effort, turn_detection {type: server_vad, threshold 0.85, silence_duration_ms, prefix_padding_ms, idle_timeout_ms} | null, audio.input|output.format {type: audio/pcm|pcmu|pcma|opus, rate 8000–48000}, audio.*.transport json|binary, audio.input.transcription {model: grok-transcribe, language_hint, keyterms}, audio.output.speed 0.7–1.5, resumption.enabled, tools[] (function, web_search, x_search, file_search, mcp)} |
setup {model, generationConfig {responseModalities: [AUDIO], speechConfig {voiceConfig, languageCode}, mediaResolution, thinkingConfig, translationConfig, enableAffectiveDialog}, systemInstruction, tools (functionDeclarations w/ behavior, googleSearch), realtimeInputConfig {automaticActivityDetection {disabled, startOfSpeechSensitivity, endOfSpeechSensitivity, prefixPaddingMs, silenceDurationMs}, activityHandling, turnCoverage}, sessionResumption {handle}, contextWindowCompression {triggerTokens, slidingWindow{targetTokens}}, inputAudioTranscription, outputAudioTranscription, proactivity {proactiveAudio}} |
| Ephemeral auth | POST /v1/realtime/client_secrets {expires_after{seconds 10–7200}, session} → ek_… |
none (server posts the SDP) | POST /v1/realtime/client_secrets {expires_after:{seconds ≤3600, default 600}, session?} → xai-realtime… (Bearer or subprotocol xai-client-secret.<token>) — LIVE_VERIFIED |
POST /v1beta/auth_tokens {uses, expireTime (<20 h), newSessionExpireTime, bidiGenerateContentSetup | liveConnectConstraints} → auth_tokens/… (PREVIEW, LIVE_VERIFIED); token on the non-Constrained method → close 1008 |
| Events | 45 server / 11 client (session.update, input_audio_buffer.append/commit, response.create/cancel, conversation.item.create/truncate …) |
20 server / 11 client | 39 server / 9 client with the OpenAI names (session.created, conversation.created, input_audio_buffer.speech_started/stopped/committed, conversation.item.added, response.created, response.output_audio.delta/done, response.output_audio_transcript.delta/done, response.function_call_arguments.*, mcp_list_tools.*, response.mcp_call.*, response.done, error, LIVE_DISCOVERED ping) |
21 server / 11 client message types (setup → setupComplete; clientContent, realtimeInput.{audio,video,text,activityStart,activityEnd,audioStreamEnd}, toolResponse → serverContent.{modelTurn, generationComplete, turnComplete, interrupted, interactionStatus, inputTranscription, outputTranscription, groundingMetadata}, toolCall, toolCallCancellation, usageMetadata, goAway, sessionResumptionUpdate) |
| Audio formats | PCM16 / G.711 / Opus per audio.format |
per session | in/out audio/pcm 8–48 kHz (default 24 kHz), pcmu, pcma, opus; JSON base64 or binary frames |
in: 16-bit PCM 16 kHz mono (audio/pcm;rate=16000); out: 24 kHz PCM; video ≤1 fps JPEG |
| Interruption | VAD (server_vad, semantic_vad), conversation.item.truncate |
model-managed full duplex | server_vad auto-commit or manual input_audio_buffer.commit; conversation.item.truncate |
automatic activity detection (configurable sensitivities) or manual activityStart/activityEnd; serverContent.interrupted; proactive audio always on for 3.8 |
| Tools | function tools, mcp, web search via delegation |
delegated to Responses or client | function, web_search, x_search, file_search (collections), mcp inside the session |
functionDeclarations (behavior: NON_BLOCKING default on 3.8; scheduling, willContinue on responses) + googleSearch only — no Maps / URL context / code execution / file search in Live |
| Thinking | model-internal | delegated model's reasoning |
reasoning.effort: high | none (URL or session.update) |
3.8-live interleaved (omit thinkingConfig); -extended-thinking LOW/MEDIUM/HIGH + includeThoughts; 3.1-flash-live MINIMAL–HIGH; 2.5 thinkingBudget |
| Session limit | 60 minutes | billed per second; no stated cap | 120 minutes; concurrent sessions 10 / 20 / 50 / 100 / 200 by tier | connection ≈10 min (goAway {timeLeft} first); session 15 min audio / 2 min audio+video, unlimited with contextWindowCompression; resumption handle valid 2 h; transcribe-live 10 min |
| Usage / billing surface | response.done.response.usage |
per-second minutes | response.done top-level usage {…, output_audio_seconds, billable_audio_seconds} (rounded up); response.create free |
usageMetadata messages (audio ≈ 25 tokens/s; whole context re-billed per turn) |
| Legacy | POST /v1/realtime/sessions, /transcription_sessions with OpenAI-Beta: realtime=v1 → 404 |
webhook alias live.call.incoming deprecated |
grok-voice-think-fast-1.0 LEGACY |
half-cascade models (gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview) RETIRED 2025-12-09; mediaChunks, speechState deprecated; responseModalities: [TEXT] → close 1007 on every current Live model |
| Streaming twin | api: realtime, realtime-translation |
api: live |
api: realtime, tts, stt (xai) |
api: live, lyria-realtime (gemini) |
The same voice turn (text in, audio out) on the three voice stacks
// OpenAI Realtime and xAI realtime — identical client messages (xAI: wss://api.x.ai/v1/realtime?model=grok-voice-latest)
{"type": "session.update", "session": {"instructions": "Reply with OK.", "voice": "alloy"}} // xAI: "voice": "eve"
{"type": "conversation.item.create", "item": {"type": "message", "role": "user", "content": [{"type": "input_text", "text": "Reply with OK."}]}}
{"type": "response.create"}
// → response.created → response.output_item.added → response.output_audio.delta… → response.output_audio_transcript.delta… → response.done// Gemini Live — wss://…/v1beta.GenerativeService.BidiGenerateContent
{"setup": {"model": "models/gemini-3.8-live", "generationConfig": {"responseModalities": ["AUDIO"], "speechConfig": {"voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Kore"}}}}, "systemInstruction": {"parts": [{"text": "Reply with OK."}]}}}
// ← {"setupComplete": {}}
{"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Reply with OK."}]}], "turnComplete": true}}
// ← {"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "…"}}]}}} … {"serverContent": {"generationComplete": true}} {"serverContent": {"turnComplete": true}} {"usageMetadata": {…}}3. Speech APIs
| Capability | OpenAI | xAI | Gemini | Anthropic |
|---|---|---|---|---|
| Text-to-speech | POST /v1/audio/speech — gpt-4o-mini-tts ($0.60 / 1M text in, $12 / 1M audio out; instructions), tts-1 ($15 / 1M chars), tts-1-hd ($30 / 1M chars); SSE speech.audio.delta/done; Chat modalities: [text, audio] |
POST /v1/tts {text ≤60,000 chars (tags [pause] [laugh], <whisper>…), voice_id (28 voices, default eve), language (**required**, 422 otherwise), output_format {codec mp3|wav|pcm|mulaw|alaw, sample_rate, bit_rate}, speed, with_timestamps} and wss://api.x.ai/v1/tts?language=&voice=&codec= (text.delta → audio.delta); $15 / 1M characters; custom voices /v1/custom-voices (clone ≤120 s, 30 / team, create Enterprise-only); LIVE_VERIFIED |
generateContent on gemini-3.1-flash-tts-preview ($1 text in / $20 audio out per 1M ≈ $0.03 / min; batch 50 %; streaming), gemini-2.5-flash-preview-tts ($0.50 / $10), gemini-2.5-pro-preview-tts ($1 / $20) with responseModalities: ["AUDIO"] + speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName (30 voices) or multiSpeakerVoiceConfig.speakerVoiceConfigs[≤2]; raw 24 kHz PCM out (no header); free tier (10 req/day observed); PREVIEW · LIVE_VERIFIED |
— |
| Speech-to-text | POST /v1/audio/transcriptions — gpt-transcribe ($0.0045 / min), gpt-4o-transcribe(-diarize), whisper-1 ($0.006 / min; → 2027-02-26); SSE transcript.text.delta/segment/done; POST /v1/audio/translations (whisper-1 → English) |
POST /v1/stt (multipart file ≤500 MB or url; language, diarize, keyterm, vad_threshold, model: grok-voice-transcribe-2.0|1.0) $0.10 / h; wss://api.x.ai/v1/stt?sample_rate=&encoding=&interim_results=&endpointing=&smart_turn= (binary frames → transcript.created/partial/done) $0.20 / h; 25 languages; LIVE_VERIFIED (REST) |
gemini-3.5-transcribe via generateContent + generationConfig.audioTranscriptionConfig {languageCodes[], customVocabulary[] ≤1,000, mode VERBATIM|SMART, diarization, wordTimestamp} → parts[].audioTranscription {text, speakerLabel, words[]} (LIVE_DISCOVERED shape); ≤1 h audio (30 min with diarization); $2 / 1M audio in (≈ $0.003 / min) + $12 / 1M text out ≈ $0.005 / min; free tier; gemini-3.5-transcribe-live over the Live WebSocket (≈ $0.009 / min); gemini-3.5-live-translate-preview speech-to-speech translation (≈ $0.037 / min) |
— |
| Audio understanding in chat | Chat Completions input_audio on gpt-audio-1.5 ($2.5/$10 text, $32/$64 audio), gpt-audio-mini |
— | any 3.x text model: inlineData audio (WAV/MP3/AIFF/AAC/OGG/FLAC) billed as audio input tokens ($1 / 1M on 2.5 Flash / 3 Flash, $0.50 on 3.1 Flash-Lite, single price on 3.5+) |
— |
| Deprecations | whisper-1, gpt-4o-transcribe* → 2027-02-26; gpt-audio, gpt-4o-*-audio/realtime → 2027-01-20 |
grok-voice-think-fast-1.0 LEGACY; STT default-model ambiguity (release notes 1.0 vs model page 2.0) |
2.5 TTS previews → gemini-3.1-flash-tts-preview (no date); half-cascade Live models retired |
— |
4. Image generation
| OpenAI Images API / tool | xAI Imagine | Gemini image models | Anthropic | |
|---|---|---|---|---|
| Models | gpt-image-2 (2026-04-21), gpt-image-2.5-flare / -sunburst (2026-09-08); gpt-image-1, -1-mini, -1.5, chatgpt-image-latest DEPRECATED; DALL·E RETIRED |
grok-imagine-image (LIVE_VERIFIED), grok-imagine-image-2.0 (quality/resolution matrix), grok-imagine-image-quality (DEPRECATED → 2026-11-02, redirects to 2.0 low); grok-2-image RETIRED (404) |
gemini-3.1-flash-image (Nano Banana 2, GA), gemini-3.1-flash-lite-image, gemini-3-pro-image (Nano Banana Pro, GA; -preview / nano-banana-pro-preview RETIRED but listed), gemini-2.5-flash-image (DEPRECATED → 2026-10-02); Imagen 3/4 (:predict) RETIRED 2026-08-17 |
— |
| Endpoints | POST /v1/images/generations (LIVE_VERIFIED), /edits (UNVERIFIED), /variations RETIRED; tools[type=image_generation] on Responses |
POST /v1/images/generations and POST /v1/images/edits (JSON only: image {url|file_id} or images[] 2–5; OpenAI images.edit() multipart unsupported) — both LIVE_VERIFIED, sync ~5–10 s; tools[type=image_generation] on Responses; Batch (standard rates); gRPC Image/GenerateImage |
POST /v1beta/models/{image-model}:generateContent with generationConfig.responseModalities: ["TEXT","IMAGE"]; editing = image input + multi-turn; OpenAI-compat POST /v1beta/openai/images/generations (subset); Interactions response_format {type: image}; batch 50 % |
— |
| Parameters | prompt, size (any WxH ≤4K or auto), quality low…max/auto, background, output_format, output_compression, moderation, stream, partial_images, n; edits: 1–16 images, mask, input_fidelity |
prompt, n 1–10, response_format url|b64_json, aspect_ratio (incl. 21:9, 5:2, auto), resolution 1k|1.5k|2k, quality low|medium|auto (2.0), storage_options {filename, expires_after, public_url}, user; no size/mask/seed; always JPEG |
imageConfig {aspectRatio: 1:1 … 21:9 (14 ratios), imageSize: 512|1K|2K|4K}, thinkingConfig.thinkingLevel MINIMAL|HIGH (3.1 image models); Google Search grounding on 3.1-flash-image / 3-pro-image; SynthID (+ C2PA on Lite) |
— |
| Price | text in $5, image in $8, image out $30 per 1M tokens (≈ $0.006 low 1024² → $0.211 high 1024²); batch 50 % on gpt-image-2 | per image: $0.02 (image), $0.04 (2.0 low/1k) … $0.08 (2.0 medium/2k), $0.05 (quality); edits bill input + output image (live $0.022) | per image equivalents: 3.1-flash-image $0.045 (0.5K) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) (image out $60 / 1M); 3.1-flash-lite-image $0.034 (1K); 3-pro-image $0.134 (1K/2K) / $0.24 (4K) ($120 / 1M, priority $216); 2.5-flash-image $0.039 | — |
| Streaming | image_generation.partial_image / .completed; tool events response.image_generation_call.* |
none (sync) | none (unary generateContent; streaming returns the image part in a chunk) |
— |
| Status / access | LIVE_VERIFIED | LIVE_VERIFIED; not served on us.api.x.ai; rate limit 6 → 100 RPS by tier |
ACCOUNT_RESTRICTED (paid tier only, limit: 0 on free) |
— |
5. Video generation
| OpenAI (deprecated) | xAI Imagine video | Gemini Veo 3.1 / Omni | Anthropic | |
|---|---|---|---|---|
| Models / price | sora-2 $0.10 / s (720p), sora-2-pro $0.30–$0.70 / s — Videos API + models shut down 2026-09-24 |
grok-imagine-video $0.05 / s (LIVE_VERIFIED: 1 s = $0.05), grok-imagine-video-1.5 $0.08 / s (native 1080p, reference-to-video, reference_audios, last_frame; aliases -preview, -2026-05-30) |
veo-3.1-generate-preview $0.40 / s (720p/1080p), $0.60 / s (4K); veo-3.1-fast-generate-preview $0.10 / $0.12 / $0.30; veo-3.1-lite-generate-preview $0.05 / $0.08 (no 4K); charged only on success; gemini-omni-1.1-flash (Interactions; video out $17.50 / 1M tokens ≈ $0.10 / s at 720p); Veo 2.0/3.0 RETIRED 2026-06-30 |
— |
| Endpoints | POST /v1/videos, GET /v1/videos[/{id}[/content]], /remix, /edits, /extensions, /characters; webhooks video.completed/failed |
POST /v1/videos/generations {model, prompt, duration 1–15 (default 8), aspect_ratio, resolution 480p|720p|1080p, generate_audio, image {url|file_id}, reference_images[], reference_audios[], last_frame, storage_options} → {request_id}; GET /v1/videos/{request_id} → 202 {status: pending, progress} → 200 {status: done, video:{url, duration, respect_moderation}, usage} | failed {error} | expired; POST /v1/videos/edits (video ≤8.7 s, 720p cap), POST /v1/videos/extensions (+2–10 s); Batch (URLs expire 1 h); gRPC Video/* |
POST /v1beta/models/veo-3.1-*:predictLongRunning {instances[{prompt, image, lastFrame, referenceImages[≤3]{image, referenceType}, video}], parameters {aspectRatio 16:9|9:16, resolution 720p|1080p|4k, durationSeconds 4|6|8, personGeneration, negativePrompt, seed, numberOfVideos}} → Operation {name}; poll GET /v1beta/{name} (11 s–6 min); download files/{id}:download?alt=media (stored 2 days); extension +7 s ≤20×; OpenAI-compat POST /v1beta/openai/videos |
— |
| Status | DEPRECATED · LIVE_VERIFIED | LIVE_VERIFIED (generation + poll); edits 422 / extensions not tested; rate limit 10 → 158 RPS | PREVIEW; predictLongRunning 400 on this free-tier key; lite model GET LIVE_VERIFIED |
— |
6. Music, embeddings and moderation
| OpenAI | xAI | Gemini | Anthropic | |
|---|---|---|---|---|
| Music | — | — | lyria-3.5 (GA 2026-09-03, $0.08 / song, MP3; WAV via Interactions response_format {type: audio}), lyria-3-clip-preview ($0.04 / 30-s clip), lyria-3-pro-preview ($0.08); Lyria RealTime wss://…/v1beta.GenerativeService.BidiGenerateMusic (setup {model: models/lyria-realtime-exp} → setupComplete; clientContent.weightedPrompts[], musicGenerationConfig {bpm 60–200, density, brightness, scale, guidance, musicGenerationMode, temperature, seed}, playbackControl PLAY|PAUSE|STOP|RESET_CONTEXT → serverContent.audioChunks[] {data, mimeType: audio/l16;rate=48000;channels=2}, filteredPrompt) — experimental, unpriced, LIVE_VERIFIED; 3.x models ACCOUNT_RESTRICTED (no free tier) |
— |
| Embeddings | POST /v1/embeddings — text-embedding-3-small (1536 dims, $0.02 / 1M), -3-large (3072, $0.13), ada-002 (LEGACY); dimensions, encoding_format; 8,192 tokens / input, 2,048 inputs |
POST /v1/embeddings {model, input, encoding_format, dimensions} documented (OpenAPI) — grok-embedding-small → 404 'does not exist or your team does not have access', GET /v1/embedding-models → {models: []}, no published price → ACCOUNT_RESTRICTED; used internally by Collections; gRPC Embedder.Embed |
POST /v1beta/models/gemini-embedding-2:embedContent {content, outputDimensionality 128–3072, embedContentConfig {autoTruncate, documentOcr, audioTrackExtraction}} / :batchEmbedContents {requests[]} — multimodal (text, ≤6 images, ≤180 s audio, ≤120 s video, PDF ≤6 pages → one vector), 8,192 tokens, no taskType (use task: / query: prefixes); $0.20 text / $0.45 image ($0.00012 / image) / $6.50 audio ($0.00016 / s) / $12 video ($0.00079 / frame) per 1M; batch 50 %; free tier; gemini-embedding-001 (text, 2,048 tokens, 8 taskType values) DEPRECATED → 2028-05-14; :asyncBatchEmbedContent ACCOUNT_RESTRICTED; OpenAI-compat /v1beta/openai/embeddings; LIVE_VERIFIED |
— (use a third-party embedder; results go back as search_result blocks) |
| Moderation | POST /v1/moderations (omni-moderation-latest, text + image, 13 categories, free); inline moderation {model, policy}; text-moderation-* RETIRED |
none; usage-guideline violations are billed (+ $0.05 pre-generation fee on Responses); respect_moderation on media results; failed.error.code on videos |
none; safetySettings[] {category, threshold} (default Off on 2.5/3.x), safetyRatings[], promptFeedback.blockReason, finishReason SAFETY|RECITATION|SPII|IMAGE_SAFETY…; enablePromptInjectionDetection on computer use |
none (stop_reason: refusal + stop_details.category; beta fallbacks) |
7. Vision and documents — where Anthropic is at parity or ahead
| Topic | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Image sources | URL, data URL, file_id |
base64, URL (fetched by Anthropic), file_id |
https URL or data URL (input_image.image_url; file_id variant → 400); Chat image_url {url, detail}; Imagine edits accept file_id |
inlineData base64, fileData {fileUri} (Files API 2 GB / 48 h, GCS via files:register), public URLs via urlContext |
| Formats | PNG, JPEG, WEBP, non-animated GIF | JPEG, PNG, GIF (first frame), WEBP | JPEG, PNG (≥512 px) | PNG, JPEG, WEBP, HEIC, HEIF |
| Limits | ≤1,500 images, ≤512 MB per request | 100 images (200k models) / 600 (1M models); 10 MB per image | not documented per request; images billed as text-price tokens (image_input = input) |
many images per request (token-bounded, 1M context); mediaResolution LOW/MEDIUM/HIGH (+ per-part ULTRA_HIGH on Gemini 3) |
| Token cost | patch models ⌈w/32⌉·⌈h/32⌉ × multiplier; tile models base + 170 per tile; detail |
⌈w/28⌉ × ⌈h/28⌉ after downscale; hi-res 2576 px / 4,784 tokens (4.7+) | usage.prompt_tokens_details.image_tokens; same $/token as text |
docs 280 / 560 / 1,120 / 2,240 tokens by resolution level (observed 256 / 529 / 1,089 / 2,209); promptTokensDetails[] {modality: IMAGE} |
input_file (< 50 MB); text + page images |
document block; ≤600 pages; 32 MB body; page_location citations |
input_file → implicit attachment search ($10 / 1k calls) |
application/pdf inline or Files; pages billed at image rate (DOCUMENT modality, 560 tokens / page in countTokens) |
|
| Audio / video | Chat audio models only | — | — | native audio and video parts on every 3.x text model; videoMetadata {startOffset, endOffset, fps}; YouTube URLs |
| Citations | annotations only for hosted-tool results | citations {enabled: true} on any document / search_result → char/page/block-level |
url_citation annotations for search tools; collections citations collections://… |
groundingMetadata / urlContextMetadata for tool results; citationMetadata (recitation) — no citations for your own documents |
| Count before sending | POST /v1/responses/input_tokens (images/files) |
POST /v1/messages/count_tokens (live: 1×1 PNG + "Describe." = 15 tokens) |
POST /v1/tokenize-text (text only) |
POST /v1beta/models/{model}:countTokens (images, PDF, audio, video, cached content) |
The same image question on all four providers
// OpenAI
{"model": "gpt-5.4-nano", "max_output_tokens": 64,
"input": [{"role": "user", "content": [{"type": "input_text", "text": "What is in this image?"}, {"type": "input_image", "image_url": "https://example.com/cat.png", "detail": "low"}]}]}// Anthropic
{"model": "claude-haiku-4-5-20251001", "max_tokens": 64,
"messages": [{"role": "user", "content": [{"type": "text", "text": "What is in this image?"}, {"type": "image", "source": {"type": "url", "url": "https://example.com/cat.png"}}]}]}// xAI (same item shape as OpenAI; reasoning tokens billed)
{"model": "grok-4.3", "max_output_tokens": 64, "reasoning": {"effort": "low"},
"input": [{"role": "user", "content": [{"type": "input_text", "text": "What is in this image?"}, {"type": "input_image", "image_url": "https://example.com/cat.png", "detail": "low"}]}]}// Gemini (inline bytes or a Files API URI; no public-URL fetch except via urlContext)
{"contents": [{"role": "user", "parts": [{"text": "What is in this image?"}, {"inlineData": {"mimeType": "image/png", "data": "<base64>"}}]}],
"generationConfig": {"maxOutputTokens": 64, "mediaResolution": "MEDIA_RESOLUTION_LOW"}}Related: features · streaming · pricing · docs/openai/realtime.md · docs/openai/live.md · docs/openai/audio.md · docs/openai/images.md · docs/openai/video.md · docs/openai/embeddings.md · docs/openai/moderation.md · docs/xai/voice.md · docs/xai/images.md · docs/xai/videos.md · docs/gemini/live-api.md · docs/gemini/speech-generation.md · docs/gemini/transcription.md · docs/gemini/image-generation.md · docs/gemini/video-generation.md · docs/gemini/music-generation.md · docs/gemini/embeddings.md · docs/anthropic/vision-and-documents.md · docs/anthropic/citations.md.