# Schema validation of model output **Status:** DOCUMENTED (structured-output parameters verified by the provider agents' tests in `tests/{openai,anthropic,xai,gemini}`; xAI chat `response_format` + Responses `text.format` and Gemini `responseJsonSchema`/`responseFormat` LIVE_VERIFIED 2026-09-18/19) **Sources:** OpenAI OpenAPI spec (`TextResponseFormatConfiguration`, function `strict`) · https://developers.openai.com/api/docs/guides/structured-outputs · https://platform.claude.com/docs/en/build-with-claude/structured-outputs · https://platform.claude.com/docs/en/api/handling-stop-reasons · xAI: https://docs.x.ai/developers/model-capabilities/text/structured-outputs (`response_format.json_schema {name, strict, schema}` / `text.format`; `additionalProperties` defaults **false**; enforced vs best-effort keywords; rejected constructs; ECMA-262 `pattern` subset; tool schemas "strict flag implicitly true") · Gemini: https://ai.google.dev/gemini-api/docs/structured-output (`responseMimeType` + `responseJsonSchema`, new `responseFormat.text {mimeType, schema}`, `propertyOrdering`, supported keywords, "values are not validated semantically — validate client-side"), https://ai.google.dev/api/generate-content (#GenerationConfig, #Schema, #FinishReason `MAX_TOKENS`, `MALFORMED_FUNCTION_CALL`), https://ai.google.dev/gemini-api/docs/function-calling (`mode: VALIDATED`) **Last verified:** 2026-09-19 ## Why validate twice Provider-side structured outputs constrain *generation*; your code must still validate before *acting*, because: - `strict` guarantees the shape, not the meaning (`{"amount": 1e12}` is a valid number; `{"path": "../.."}` is a valid string). - Output can be truncated (`stop_reason: max_tokens` / `status: incomplete`) → invalid JSON. - The model can refuse (`refusal` content part / `stop_reason: refusal`) → no JSON at all. - Provider limitations mean some schema keywords are ignored (both docs list unsupported JSON Schema features) — a `pattern` or `minimum` you wrote may not have been enforced. - Anything not in `strict` mode is best-effort. ## Provider parameters | | OpenAI Responses | Anthropic Messages | xAI (Responses / Chat) | Gemini generateContent | |---|---|---|---|---| | JSON output | `text.format = {type:"json_schema", name, schema, strict: true}`; `json_object` | `output_config.format = {type:"json_schema", schema}` | `text.format` (Responses) / `response_format = {type:"json_schema", json_schema:{name, strict, schema}}` (Chat); `json_object`; `additionalProperties` **defaults to false** (set `true` explicitly to allow extras); non-`required` fields optional | `generationConfig.responseMimeType: "application/json"` + `responseJsonSchema` (works even without the mime type — not enforced), or `responseFormat.text {mimeType: APPLICATION_JSON, schema}` (new; wire enum is UPPERCASE), or legacy `responseSchema` (OpenAPI subset, `propertyOrdering`); `text/x.enum` for enums | | Tool arguments | `tools[].strict: true` (`additionalProperties: false`, all `required`) | `tools[].strict: true` | tool `parameters` / `input_schema` **always strictly enforced** ("strict flag implicitly true"; the field is accepted and ignored) | `toolConfig.functionCallingConfig.mode: VALIDATED` validates calls against declarations; otherwise best-effort — `finishReason: MALFORMED_FUNCTION_CALL` when the model's call does not parse | | Refusal signal | `{type:"refusal"}` part; `response.refusal.delta` | `stop_reason: "refusal"` | Responses: refusal part (OpenAI shape); Chat: `finish_reason: "content_filter"`; policy violations are billed | `promptFeedback.blockReason` (no candidates) or `finishReason: SAFETY \| PROHIBITED_CONTENT \| SPII \| BLOCKLIST \| RECITATION` — **HTTP 200** in all cases | | Truncation signal | `status: "incomplete"` + `incomplete_details.reason: "max_output_tokens"` | `stop_reason: "max_tokens"` | Responses: `incomplete_details.reason ∈ max_output_tokens \| max_prompt_tokens \| max_time_limit`; Chat: `finish_reason: "length"` | `finishReason: "MAX_TOKENS"` — JSON cut mid-way (`{"ok": true, "word": "OK` observed with 32 tokens; thinking tokens count against the same cap) | | Limitations | JSON Schema subset, root must be object | "JSON Schema limitations" | Enforced: types, enum/const, anyOf/oneOf, single-subschema allOf, `$ref/$defs` (non-circular), `format` date/time/date-time/email/uuid/ipv4/ipv6/uri, min/max, length ≤ 2048, items ≤ 256, properties ≤ 64. Best-effort: `not`, `if/then/else`, multi-`allOf`, unknown formats. **400**: empty `enum`/`anyOf`, boolean property schemas, `min/maxContains`, `items` as array. `pattern` is an ECMA-262 subset (no backreferences, lookaround, `\b`; implicit anchors) | Supported: `$id/$defs/$ref/$anchor`, `type` (arrays incl. `"null"`), `format` date/time/date-time, `enum`, `items`, `prefixItems`, `minItems/maxItems`, `minimum/maximum`, `anyOf` (`oneOf` treated as anyOf), `properties`, `additionalProperties`, `required`, `propertyOrdering`. **Unsupported keywords are silently ignored** (`minLength`, `pattern` in JSON-Schema form); cyclic `$ref` only in non-required properties; very large schemas rejected; legacy `responseSchema` 400s on unknown keys; `responseSchema` + `responseJsonSchema` together → both ignored, prose returned | The shared adapters (`examples/shared/provider-abstraction/`) set `strict: true` on OpenAI/xAI by default, map `json_schema` to each provider's parameters, and raise on refusal (and on Gemini `MAX_TOKENS`) before `json.loads`. Because xAI defaults `additionalProperties` to false and Gemini ignores it in some forms, a schema that "passes" on one provider may reject or over-accept on another — generate per-provider variants from one source. ## Validation pipeline (your code) 1. **Check stop reason first** (`end` vs `max_tokens`/`incomplete`/`refusal`); never parse a truncated body. 2. **Parse strictly** (`json.loads` / `JSON.parse`); reject trailing garbage, NaN/Infinity, duplicate keys if your parser allows. 3. **Validate against your schema again** with a full validator (`jsonschema` Draft 2020-12 in Python, Ajv in TS) — including the keywords the provider may ignore (`pattern`, `minimum`, `maxLength`, `enum`, `format`). 4. **Semantic checks**: ranges, referential integrity (ids exist and belong to this user), path confinement, URL allowlists, currency/amount limits. 5. **Type-safe binding**: pydantic / zod models with `extra = forbid` → the rest of your code never touches raw dicts. 6. **Fail closed**: on any failure, do not act; either re-prompt with the validation error (bounded retries) or escalate. Log the raw output with the request id. 7. **Schema hygiene**: `additionalProperties: false` everywhere; enums over free strings; small integer ranges; explicit `null` handling; descriptions that tell the model *what* not *how*. Keep one schema source and generate both provider variants (they diverge in supported keywords). ## Streaming Structured output can be streamed (`response.output_text.delta` / `text_delta`), and tool arguments always are. Use the partial-JSON preview only for UI; validate the *final* string (`docs/architecture/streaming-patterns.md`). ## Checklist - [ ] `strict: true` (OpenAI, xAI) / `strict` tools (Anthropic) / `mode: VALIDATED` (Gemini) wherever supported; `additionalProperties: false` explicit everywhere (xAI defaults it, Gemini may ignore it). - [ ] Stop reason checked before parsing; refusal and truncation handled explicitly — including Gemini's HTTP-200 blocks (`promptFeedback`, `finishReason`) and xAI `max_time_limit`. - [ ] Keywords the provider ignores (Gemini `minLength`/`pattern`, xAI best-effort `not`/`if`) re-validated in code. - [ ] Full JSON Schema validation + semantic checks in code; typed models with unknown-field rejection. - [ ] Fail closed with bounded re-prompting; raw output logged with request id. - [ ] One schema source → per-provider variants tested against each provider's limitations. - [ ] Never act on streamed partial arguments.