Schema validation of model output
Status: DOCUMENTED (structured-output parameters verified by the provider agents' tests in tests/{openai,anthropic,xai,gemini}; xAI chat response_format + Responses text.format and Gemini responseJsonSchema/responseFormat LIVE_VERIFIED 2026-09-18/19)
Sources: OpenAI OpenAPI spec (TextResponseFormatConfiguration, function strict) · https://developers.openai.com/api/docs/guides/structured-outputs · https://platform.claude.com/docs/en/build-with-claude/structured-outputs · https://platform.claude.com/docs/en/api/handling-stop-reasons · xAI: https://docs.x.ai/developers/model-capabilities/text/structured-outputs (response_format.json_schema {name, strict, schema} / text.format; additionalProperties defaults false; enforced vs best-effort keywords; rejected constructs; ECMA-262 pattern subset; tool schemas "strict flag implicitly true") · Gemini: https://ai.google.dev/gemini-api/docs/structured-output (responseMimeType + responseJsonSchema, new responseFormat.text {mimeType, schema}, propertyOrdering, supported keywords, "values are not validated semantically — validate client-side"), https://ai.google.dev/api/generate-content (#GenerationConfig, #Schema, #FinishReason MAX_TOKENS, MALFORMED_FUNCTION_CALL), https://ai.google.dev/gemini-api/docs/function-calling (mode: VALIDATED)
Last verified: 2026-09-19
Why validate twice
Provider-side structured outputs constrain generation; your code must still validate before acting, because:
strictguarantees the shape, not the meaning ({"amount": 1e12}is a valid number;{"path": "../.."}is a valid string).- Output can be truncated (
stop_reason: max_tokens/status: incomplete) → invalid JSON. - The model can refuse (
refusalcontent part /stop_reason: refusal) → no JSON at all. - Provider limitations mean some schema keywords are ignored (both docs list unsupported JSON Schema features) — a
patternorminimumyou wrote may not have been enforced. - Anything not in
strictmode is best-effort.
Provider parameters
| OpenAI Responses | Anthropic Messages | xAI (Responses / Chat) | Gemini generateContent | |
|---|---|---|---|---|
| JSON output | text.format = {type:"json_schema", name, schema, strict: true}; json_object |
output_config.format = {type:"json_schema", schema} |
text.format (Responses) / response_format = {type:"json_schema", json_schema:{name, strict, schema}} (Chat); json_object; additionalProperties defaults to false (set true explicitly to allow extras); non-required fields optional |
generationConfig.responseMimeType: "application/json" + responseJsonSchema (works even without the mime type — not enforced), or responseFormat.text {mimeType: APPLICATION_JSON, schema} (new; wire enum is UPPERCASE), or legacy responseSchema (OpenAPI subset, propertyOrdering); text/x.enum for enums |
| Tool arguments | tools[].strict: true (additionalProperties: false, all required) |
tools[].strict: true |
tool parameters / input_schema always strictly enforced ("strict flag implicitly true"; the field is accepted and ignored) |
toolConfig.functionCallingConfig.mode: VALIDATED validates calls against declarations; otherwise best-effort — finishReason: MALFORMED_FUNCTION_CALL when the model's call does not parse |
| Refusal signal | {type:"refusal"} part; response.refusal.delta |
stop_reason: "refusal" |
Responses: refusal part (OpenAI shape); Chat: finish_reason: "content_filter"; policy violations are billed |
promptFeedback.blockReason (no candidates) or finishReason: SAFETY | PROHIBITED_CONTENT | SPII | BLOCKLIST | RECITATION — HTTP 200 in all cases |
| Truncation signal | status: "incomplete" + incomplete_details.reason: "max_output_tokens" |
stop_reason: "max_tokens" |
Responses: incomplete_details.reason ∈ max_output_tokens | max_prompt_tokens | max_time_limit; Chat: finish_reason: "length" |
finishReason: "MAX_TOKENS" — JSON cut mid-way ({"ok": true, "word": "OK observed with 32 tokens; thinking tokens count against the same cap) |
| Limitations | JSON Schema subset, root must be object | "JSON Schema limitations" | Enforced: types, enum/const, anyOf/oneOf, single-subschema allOf, $ref/$defs (non-circular), format date/time/date-time/email/uuid/ipv4/ipv6/uri, min/max, length ≤ 2048, items ≤ 256, properties ≤ 64. Best-effort: not, if/then/else, multi-allOf, unknown formats. 400: empty enum/anyOf, boolean property schemas, min/maxContains, items as array. pattern is an ECMA-262 subset (no backreferences, lookaround, \b; implicit anchors) |
Supported: $id/$defs/$ref/$anchor, type (arrays incl. "null"), format date/time/date-time, enum, items, prefixItems, minItems/maxItems, minimum/maximum, anyOf (oneOf treated as anyOf), properties, additionalProperties, required, propertyOrdering. Unsupported keywords are silently ignored (minLength, pattern in JSON-Schema form); cyclic $ref only in non-required properties; very large schemas rejected; legacy responseSchema 400s on unknown keys; responseSchema + responseJsonSchema together → both ignored, prose returned |
The shared adapters (examples/shared/provider-abstraction/) set strict: true on OpenAI/xAI by default, map json_schema to each provider's parameters, and raise on refusal (and on Gemini MAX_TOKENS) before json.loads. Because xAI defaults additionalProperties to false and Gemini ignores it in some forms, a schema that "passes" on one provider may reject or over-accept on another — generate per-provider variants from one source.
Validation pipeline (your code)
- Check stop reason first (
endvsmax_tokens/incomplete/refusal); never parse a truncated body. - Parse strictly (
json.loads/JSON.parse); reject trailing garbage, NaN/Infinity, duplicate keys if your parser allows. - Validate against your schema again with a full validator (
jsonschemaDraft 2020-12 in Python, Ajv in TS) — including the keywords the provider may ignore (pattern,minimum,maxLength,enum,format). - Semantic checks: ranges, referential integrity (ids exist and belong to this user), path confinement, URL allowlists, currency/amount limits.
- Type-safe binding: pydantic / zod models with
extra = forbid→ the rest of your code never touches raw dicts. - Fail closed: on any failure, do not act; either re-prompt with the validation error (bounded retries) or escalate. Log the raw output with the request id.
- Schema hygiene:
additionalProperties: falseeverywhere; enums over free strings; small integer ranges; explicitnullhandling; descriptions that tell the model what not how. Keep one schema source and generate both provider variants (they diverge in supported keywords).
Streaming
Structured output can be streamed (response.output_text.delta / text_delta), and tool arguments always are. Use the partial-JSON preview only for UI; validate the final string (docs/architecture/streaming-patterns.md).
Checklist
-
strict: true(OpenAI, xAI) /stricttools (Anthropic) /mode: VALIDATED(Gemini) wherever supported;additionalProperties: falseexplicit everywhere (xAI defaults it, Gemini may ignore it). - Stop reason checked before parsing; refusal and truncation handled explicitly — including Gemini's HTTP-200 blocks (
promptFeedback,finishReason) and xAImax_time_limit. - Keywords the provider ignores (Gemini
minLength/pattern, xAI best-effortnot/if) re-validated in code. - Full JSON Schema validation + semantic checks in code; typed models with unknown-field rejection.
- Fail closed with bounded re-prompting; raw output logged with request id.
- One schema source → per-provider variants tested against each provider's limitations.
- Never act on streamed partial arguments.