OpenAI Embeddings API (POST /v1/embeddings)
Status: DOCUMENTED + LIVE_VERIFIED (all three models, dimensions, encoding_format=base64, token-array input, error shapes).
Sources: Embeddings reference · Vector embeddings guide · Pricing · model pages text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002 · OpenAPI CreateEmbeddingRequest / CreateEmbeddingResponse.
Last verified: 2026-09-18.
Machine-readable: endpoints fragment, parameters/openai-embeddings.json, objects/openai-media-objects.json, prices/openai-media.json.
Model matrix
| Model | Native dims | dimensions param |
Max input tokens | Price (per 1M input tokens) | ~Pages / $ | MTEB | Knowledge | Batch | Live (input "OK") |
|---|---|---|---|---|---|---|---|---|---|
text-embedding-3-small |
1536 | ✓ (1…1536) | 8 192 | $0.02 | 62 500 | 62.3 % | ≤ Sep 2021 | ✓ | 200 · 1536 floats · prompt_tokens: 1 · ‖v‖₂ = 1.0001 |
text-embedding-3-large |
3072 | ✓ (1…3072) | 8 192 | $0.13 | 9 615 | 64.6 % | ≤ Sep 2021 | ✓ | 200 · 3072 floats · ‖v‖₂ = 0.9998 |
text-embedding-ada-002 |
1536 | ✗ (400) | 8 192 | $0.10 | 12 500 | 61.0 % | — | ✓ | 200 · 1536 floats · response model: "text-embedding-ada-002-v2" · ‖v‖₂ = 1.0000 |
Standard-tier prices; Batch tier prices for embeddings are not broken out on the pricing page (Batch supported per model pages). Regional endpoints: /v1/embeddings available in all regions; UAE lists text-embedding-3-large specifically. Observed header on our key: x-ratelimit-limit-requests: 10000 (account-specific, not a documented limit).
Request
| Param | Type | Required | Default | Notes |
|---|---|---|---|---|
input |
string | string[] | int[] | int[][] |
✓ | — | non-empty; ≤ 8 192 tokens per input; arrays 1–2 048 items; ≤ 300 000 tokens summed per request; tokenizer cl100k_base (tiktoken) |
model |
string | ✓ | — | see matrix |
dimensions |
int | — | native | text-embedding-3-* only (Matryoshka); API returns re-normalized vectors |
encoding_format |
float | base64 |
— | float |
base64 = raw little-endian float32 bytes (4 × dims), smaller/faster to parse |
user |
string | — | — | end-user id for abuse monitoring |
Live probes: dimensions: 16 + base64 on 3-small → embedding = 88 base64 chars = 64 bytes = 16 float32, L2 norm 1.0003; dimensions on ada-002 → 400 invalid_request_error "This model does not support specifying dimensions." (param: null, plus a non-standard top-level detail object); input: [[11380]] (token array) → 200, prompt_tokens: 1.
Response
{"object":"list","model":"text-embedding-3-small",
"data":[{"object":"embedding","index":0,"embedding":[/* 1536 floats or a base64 string */]}],
"usage":{"prompt_tokens":1,"total_tokens":1}}data[i].index matches the position in the input array. Billing = usage.total_tokens × price.
Normalization, distance, dimensions
- OpenAI embeddings are unit-length (L2 = 1) — cosine similarity = dot product; cosine and Euclidean rankings are identical (guide FAQ; confirmed live to ±3e-4).
- Prefer the
dimensionsparameter over manual truncation. If you truncate client-side, re-normalize (guide showsnormalize_l2).text-embedding-3-largecut to 256 dims still beats full ada-002 on MTEB (guide). - Use cases in the guide: search, clustering, recommendations, anomaly detection, classification, zero-shot classification, 2-D t-SNE visualisation, features for regression/classification, cold-start recommendations, code search.
Batching and limits
- Up to 2 048 inputs per call and 300 k tokens total; count tokens first with tiktoken (
cl100k_base). - For large offline jobs use the Batch API (
/v1/embeddingssupported by all three models) — 24 h window, discounted tier. - Empty strings are rejected; token arrays let you pre-tokenize and control truncation yourself.
Examples and tests
examples/openai/embeddings/ — create.sh|py|ts (LIVE_VERIFIED; the .py also shows dimensions + base64 decoding). tests/openai/test_embeddings.py (cheap, always runs).