SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.3 KB

# OpenAI Embeddings API (POST /v1/embeddings)

Status: DOCUMENTED + LIVE_VERIFIED (all three models, dimensions, encoding_format=base64, token-array input, error shapes). Sources: Embeddings reference · Vector embeddings guide · Pricing · model pages text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002 · OpenAPI CreateEmbeddingRequest / CreateEmbeddingResponse. Last verified: 2026-09-18. Machine-readable: endpoints fragment, parameters/openai-embeddings.json, objects/openai-media-objects.json, prices/openai-media.json.

# Model matrix

Model Native dims dimensions param Max input tokens Price (per 1M input tokens) ~Pages / $ MTEB Knowledge Batch Live (input "OK")
text-embedding-3-small 1536 ✓ (1…1536) 8 192 $0.02 62 500 62.3 % ≤ Sep 2021 ✓ 200 · 1536 floats · prompt_tokens: 1 · ‖v‖₂ = 1.0001
text-embedding-3-large 3072 ✓ (1…3072) 8 192 $0.13 9 615 64.6 % ≤ Sep 2021 ✓ 200 · 3072 floats · ‖v‖₂ = 0.9998
text-embedding-ada-002 1536 ✗ (400) 8 192 $0.10 12 500 61.0 % — ✓ 200 · 1536 floats · response model: "text-embedding-ada-002-v2" · ‖v‖₂ = 1.0000

Standard-tier prices; Batch tier prices for embeddings are not broken out on the pricing page (Batch supported per model pages). Regional endpoints: /v1/embeddings available in all regions; UAE lists text-embedding-3-large specifically. Observed header on our key: x-ratelimit-limit-requests: 10000 (account-specific, not a documented limit).

# Request

Param Type Required Default Notes
input string | string[] | int[] | int[][] ✓ — non-empty; ≤ 8 192 tokens per input; arrays 1–2 048 items; ≤ 300 000 tokens summed per request; tokenizer cl100k_base (tiktoken)
model string ✓ — see matrix
dimensions int — native text-embedding-3-* only (Matryoshka); API returns re-normalized vectors
encoding_format float | base64 — float base64 = raw little-endian float32 bytes (4 × dims), smaller/faster to parse
user string — — end-user id for abuse monitoring

Live probes: dimensions: 16 + base64 on 3-small → embedding = 88 base64 chars = 64 bytes = 16 float32, L2 norm 1.0003; dimensions on ada-002 → 400 invalid_request_error "This model does not support specifying dimensions." (param: null, plus a non-standard top-level detail object); input: [[11380]] (token array) → 200, prompt_tokens: 1.

# Response

json
{"object":"list","model":"text-embedding-3-small",
 "data":[{"object":"embedding","index":0,"embedding":[/* 1536 floats or a base64 string */]}],
 "usage":{"prompt_tokens":1,"total_tokens":1}}

data[i].index matches the position in the input array. Billing = usage.total_tokens × price.

# Normalization, distance, dimensions

  • OpenAI embeddings are unit-length (L2 = 1) — cosine similarity = dot product; cosine and Euclidean rankings are identical (guide FAQ; confirmed live to ±3e-4).
  • Prefer the dimensions parameter over manual truncation. If you truncate client-side, re-normalize (guide shows normalize_l2). text-embedding-3-large cut to 256 dims still beats full ada-002 on MTEB (guide).
  • Use cases in the guide: search, clustering, recommendations, anomaly detection, classification, zero-shot classification, 2-D t-SNE visualisation, features for regression/classification, cold-start recommendations, code search.

# Batching and limits

  • Up to 2 048 inputs per call and 300 k tokens total; count tokens first with tiktoken (cl100k_base).
  • For large offline jobs use the Batch API (/v1/embeddings supported by all three models) — 24 h window, discounted tier.
  • Empty strings are rejected; token arrays let you pre-tokenize and control truncation yourself.

# Examples and tests

examples/openai/embeddings/ — create.sh|py|ts (LIVE_VERIFIED; the .py also shows dimensions + base64 decoding). tests/openai/test_embeddings.py (cheap, always runs).