SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.9 KB

# Schema validation of model output

Status: DOCUMENTED (structured-output parameters verified by the provider agents' tests in tests/{openai,anthropic,xai,gemini}; xAI chat response_format + Responses text.format and Gemini responseJsonSchema/responseFormat LIVE_VERIFIED 2026-09-18/19) Sources: OpenAI OpenAPI spec (TextResponseFormatConfiguration, function strict) · https://developers.openai.com/api/docs/guides/structured-outputs · https://platform.claude.com/docs/en/build-with-claude/structured-outputs · https://platform.claude.com/docs/en/api/handling-stop-reasons · xAI: https://docs.x.ai/developers/model-capabilities/text/structured-outputs (response_format.json_schema {name, strict, schema} / text.format; additionalProperties defaults false; enforced vs best-effort keywords; rejected constructs; ECMA-262 pattern subset; tool schemas "strict flag implicitly true") · Gemini: https://ai.google.dev/gemini-api/docs/structured-output (responseMimeType + responseJsonSchema, new responseFormat.text {mimeType, schema}, propertyOrdering, supported keywords, "values are not validated semantically — validate client-side"), https://ai.google.dev/api/generate-content (#GenerationConfig, #Schema, #FinishReason MAX_TOKENS, MALFORMED_FUNCTION_CALL), https://ai.google.dev/gemini-api/docs/function-calling (mode: VALIDATED) Last verified: 2026-09-19

# Why validate twice

Provider-side structured outputs constrain generation; your code must still validate before acting, because:

  • strict guarantees the shape, not the meaning ({"amount": 1e12} is a valid number; {"path": "../.."} is a valid string).
  • Output can be truncated (stop_reason: max_tokens / status: incomplete) → invalid JSON.
  • The model can refuse (refusal content part / stop_reason: refusal) → no JSON at all.
  • Provider limitations mean some schema keywords are ignored (both docs list unsupported JSON Schema features) — a pattern or minimum you wrote may not have been enforced.
  • Anything not in strict mode is best-effort.

# Provider parameters

OpenAI Responses Anthropic Messages xAI (Responses / Chat) Gemini generateContent
JSON output text.format = {type:"json_schema", name, schema, strict: true}; json_object output_config.format = {type:"json_schema", schema} text.format (Responses) / response_format = {type:"json_schema", json_schema:{name, strict, schema}} (Chat); json_object; additionalProperties defaults to false (set true explicitly to allow extras); non-required fields optional generationConfig.responseMimeType: "application/json" + responseJsonSchema (works even without the mime type — not enforced), or responseFormat.text {mimeType: APPLICATION_JSON, schema} (new; wire enum is UPPERCASE), or legacy responseSchema (OpenAPI subset, propertyOrdering); text/x.enum for enums
Tool arguments tools[].strict: true (additionalProperties: false, all required) tools[].strict: true tool parameters / input_schema always strictly enforced ("strict flag implicitly true"; the field is accepted and ignored) toolConfig.functionCallingConfig.mode: VALIDATED validates calls against declarations; otherwise best-effort — finishReason: MALFORMED_FUNCTION_CALL when the model's call does not parse
Refusal signal {type:"refusal"} part; response.refusal.delta stop_reason: "refusal" Responses: refusal part (OpenAI shape); Chat: finish_reason: "content_filter"; policy violations are billed promptFeedback.blockReason (no candidates) or finishReason: SAFETY | PROHIBITED_CONTENT | SPII | BLOCKLIST | RECITATION — HTTP 200 in all cases
Truncation signal status: "incomplete" + incomplete_details.reason: "max_output_tokens" stop_reason: "max_tokens" Responses: incomplete_details.reason ∈ max_output_tokens | max_prompt_tokens | max_time_limit; Chat: finish_reason: "length" finishReason: "MAX_TOKENS" — JSON cut mid-way ({"ok": true, "word": "OK observed with 32 tokens; thinking tokens count against the same cap)
Limitations JSON Schema subset, root must be object "JSON Schema limitations" Enforced: types, enum/const, anyOf/oneOf, single-subschema allOf, $ref/$defs (non-circular), format date/time/date-time/email/uuid/ipv4/ipv6/uri, min/max, length ≤ 2048, items ≤ 256, properties ≤ 64. Best-effort: not, if/then/else, multi-allOf, unknown formats. 400: empty enum/anyOf, boolean property schemas, min/maxContains, items as array. pattern is an ECMA-262 subset (no backreferences, lookaround, \b; implicit anchors) Supported: $id/$defs/$ref/$anchor, type (arrays incl. "null"), format date/time/date-time, enum, items, prefixItems, minItems/maxItems, minimum/maximum, anyOf (oneOf treated as anyOf), properties, additionalProperties, required, propertyOrdering. Unsupported keywords are silently ignored (minLength, pattern in JSON-Schema form); cyclic $ref only in non-required properties; very large schemas rejected; legacy responseSchema 400s on unknown keys; responseSchema + responseJsonSchema together → both ignored, prose returned

The shared adapters (examples/shared/provider-abstraction/) set strict: true on OpenAI/xAI by default, map json_schema to each provider's parameters, and raise on refusal (and on Gemini MAX_TOKENS) before json.loads. Because xAI defaults additionalProperties to false and Gemini ignores it in some forms, a schema that "passes" on one provider may reject or over-accept on another — generate per-provider variants from one source.

# Validation pipeline (your code)

  1. Check stop reason first (end vs max_tokens/incomplete/refusal); never parse a truncated body.
  2. Parse strictly (json.loads / JSON.parse); reject trailing garbage, NaN/Infinity, duplicate keys if your parser allows.
  3. Validate against your schema again with a full validator (jsonschema Draft 2020-12 in Python, Ajv in TS) — including the keywords the provider may ignore (pattern, minimum, maxLength, enum, format).
  4. Semantic checks: ranges, referential integrity (ids exist and belong to this user), path confinement, URL allowlists, currency/amount limits.
  5. Type-safe binding: pydantic / zod models with extra = forbid → the rest of your code never touches raw dicts.
  6. Fail closed: on any failure, do not act; either re-prompt with the validation error (bounded retries) or escalate. Log the raw output with the request id.
  7. Schema hygiene: additionalProperties: false everywhere; enums over free strings; small integer ranges; explicit null handling; descriptions that tell the model what not how. Keep one schema source and generate both provider variants (they diverge in supported keywords).

# Streaming

Structured output can be streamed (response.output_text.delta / text_delta), and tool arguments always are. Use the partial-JSON preview only for UI; validate the final string (docs/architecture/streaming-patterns.md).

# Checklist

  • strict: true (OpenAI, xAI) / strict tools (Anthropic) / mode: VALIDATED (Gemini) wherever supported; additionalProperties: false explicit everywhere (xAI defaults it, Gemini may ignore it).
  • Stop reason checked before parsing; refusal and truncation handled explicitly — including Gemini's HTTP-200 blocks (promptFeedback, finishReason) and xAI max_time_limit.
  • Keywords the provider ignores (Gemini minLength/pattern, xAI best-effort not/if) re-validated in code.
  • Full JSON Schema validation + semantic checks in code; typed models with unknown-field rejection.
  • Fail closed with bounded re-prompting; raw output logged with request id.
  • One schema source → per-provider variants tested against each provider's limitations.
  • Never act on streamed partial arguments.