Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Secret leakage — prompts, logs, stored responses, retention23**Status:** DOCUMENTED (xAI `GET /v1/me.zdr_status` LIVE_VERIFIED `no_zdr`; xAI `GET /v1/responses/{id}` on a `store:false` response returned 200 — LIVE_DISCOVERED 2026-09-19)4**Sources:** https://developers.openai.com/api/docs/guides/background · https://developers.openai.com/api/docs/guides/tools-connectors-mcp#risks-and-safety · OpenAI OpenAPI spec · https://platform.claude.com/docs/en/manage-claude/api-and-data-retention · https://platform.claude.com/docs/en/api/errors#request-id · xAI: https://docs.x.ai/developers/faq/security (30-day encrypted retention, not used for training, team-wide ZDR and what it disables, US regional handling, HIPAA via BAA), https://docs.x.ai/developers/rest-api-reference/inference/responses (`store`, `previous_response_id`, `include: reasoning.encrypted_content`, DELETE), https://docs.x.ai/developers/advanced-api-usage/context-compaction, https://docs.x.ai/developers/files/public-urls · Gemini: https://ai.google.dev/gemini-api/terms (Unpaid vs Paid Services data use), https://ai.google.dev/gemini-api/docs/usage-policies (abuse-monitoring logs, 55 days), https://ai.google.dev/gemini-api/docs/zdr (Search/Maps grounding 30-day storage; Interactions `store`), https://ai.google.dev/gemini-api/docs/available-regions (EEA/UK/CH Paid Services only), https://ai.google.dev/gemini-api/docs/interactions (`store`, `background`, `previous_interaction_id`) · `scripts/live.py::mask` (this repo)5**Last verified:** 2026-09-1967## Where secrets leak in an LLM app89| Channel | Leak | Mitigation |10|---|---|---|11| Prompt content | keys, tokens, PII pasted into `input`/`messages` (by users, by RAG, by tool outputs such as `env` dumps) | redact before sending; never put credentials in prompts (Anthropic computer-use warns even about `<robot_credentials>`); scrub tool outputs |12| Provider storage | OpenAI stores Responses **30 days by default** (`store` defaults to `true`; also required for `previous_response_id` chaining); Anthropic stores per its retention policy unless ZDR | OpenAI `store: false` when you don't need retrieval/chaining; request **ZDR** (both providers, sales/eligibility) for regulated data; know which features are ZDR-ineligible (Anthropic publishes a feature table; OpenAI: stateful features need storage) |13| Your logs | request/response bodies, `Authorization`/`x-api-key` headers, URLs with tokens | structured logging with field allowlists; mask `sk-…`, `Bearer …`, `x-api-key` (see `mask()` regexes in `scripts/live.py`); log request ids instead of bodies |14| Model output | the model repeats secrets it saw (system prompt, tool results) to the user or into a tool argument (exfil) | keep secrets out of context in the first place; output filters for key patterns; URL allowlists on outbound tools |15| Error messages | stack traces with paths/env in `is_error` results or user-facing errors | sanitise (`untrusted-tool-outputs.md`) |16| Third parties | MCP servers, web fetch targets, webhooks | MCP: data sent is under the *server's* retention (OpenAI docs); log what you send; allowlist |17| Caches | prompt caches keyed on content; `prompt_cache_key` (OpenAI) / `cache_control` (Anthropic) | caches are per-organization and provider-internal; still avoid caching secret-bearing prefixes |18| Metadata | `metadata`, `safety_identifier`, `metadata.user_id` | use opaque/hashed ids, never emails or names (Anthropic: `user_id` must not contain identifying info) |19| Files | uploaded files readable by any key in the project/workspace | delete after use; per-tenant projects (`file-uploads-and-ssrf.md`) |2021## OpenAI knobs2223- `store: false` — response not retained for retrieval (you lose `previous_response_id`, `GET /v1/responses/{id}`, background polling beyond the short window, and dashboard logs). For chaining without storage, replay the conversation yourself.24- `include[]` — only request what you need (e.g. `reasoning.encrypted_content` lets you keep reasoning state client-side with `store: false`).25- Data controls in the org dashboard (retention, training opt-out is default for API), **Zero Data Retention** and **Data Residency** as contractual features; MCP and most tools are compatible but third-party servers are out of scope.26- `safety_identifier` should be a stable **hash** of your user id.2728## Anthropic knobs2930- **ZDR** arrangement (contact sales): no prompts/responses at rest after the response; applies to Messages and Token Counting for eligible features — check the feature eligibility table (e.g. features that inherently store data such as Files, Batches results, container reuse have their own retention). HIPAA-ready access is a separate arrangement.31- `metadata.user_id`: opaque identifier only.32- Request tracing without bodies: `request-id` header / `request_id` field in errors; `anthropic-organization-id`, `anthropic-workspace-id`.3334## xAI knobs3536- Default retention **30 days** (encrypted, not used for training, then deleted); `DELETE /v1/responses/{id}` for early removal. `store: false` is echoed, but live a `store:false` response was **still retrievable by id** — do not treat `store:false` as a privacy guarantee; the documented guarantee is ZDR.37- **Zero Data Retention** is a **team-wide console toggle** (self-serve for admins where available), visible via `GET /v1/me.zdr_status` and the `x-zero-data-retention` response header. It disables the stateful features: `store`/`previous_response_id`, Files, Collections, Batch, deferred completions, stored image/video outputs (base64 only), per-key request logging, voice history — design for **client-side state** (`include:["reasoning.encrypted_content"]` and replay `output[]`, or `POST /v1/responses/compact` blobs, which are opaque and must not be edited).38- `metadata` is **rejected** on Responses (400) — end-user attribution goes in `safety_identifier`/`user`; `prompt_cache_key`/`x-grok-conv-id` are routing hints, don't put PII in them.39- Files are **team-scoped and permanent** unless `expires_after` is set; `POST /v1/files/{id}/public-url` creates an **anonymous CDN URL** (`files-cdn.x.ai`, up to 30 days) — anyone with the URL can download; revoke with `/public-url/revoke`; deleting the file revokes it.40- US regional host (`us.api.x.ai`) keeps handling/inference/retained data in the US (not Files/Collections/tools); `eu-west-1.api.x.ai` answers but is undocumented — not a compliance control.41- MCP: your `authorization`/`headers` are forwarded to the third-party server by xAI; the server's retention applies.4243## Gemini knobs4445- **Free tier = Unpaid Services**: prompts, uploaded content and responses "may be used to provide, improve, and develop Google products", **with human review** (Terms). Never send confidential or personal data through a free-tier key; enabling a Cloud Billing account switches the project to Paid Services (not used to improve products; abuse-monitoring logs kept **55 days**). End users in the EEA/UK/CH may only be served through Paid Services.46- **No full zero-data-retention on the Developer API**: Search/Maps grounding prompts and outputs are stored 30 days and cannot be disabled; Interactions store state unless `store: false` (incompatible with `background` and chaining); Files live 48 h; File Search stores persist until deleted; explicit caches until TTL. Contractual ZDR/DPA → Vertex AI.47- `store` (per-request logging override on `generateContent`, discovery-documented; accepted live) and Interactions `store: false` reduce what AI Studio logs; **not documented** whether they exclude abuse-monitoring logs (UNVERIFIED).48- End-user attribution: `labels.safety_identifier` (Cloud-label rules: lower-case, ≤ 63 chars) — use a hash, never an e-mail.49- **Key in URL** (`?key=`) puts the secret in access logs, referrers and browser history — header only. Auth keys are bound to a service account: a leaked key = a leaked identity with whatever IAM you granted.50- `thoughtSignature` / Interactions `signature` blobs are opaque encrypted reasoning state tied to your key — store them like prompt content (they can contain the model's view of your data), and never send them to another provider.5152## This repository's discipline (reuse it)5354Keys only in `.env` (600, gitignored); `scripts/live.py` masks `sk-(ant-)?…`, `xai-…`, `AIza…`, `AQ.…`, `?key=` and `bearer|x-api-key|x-goog-api-key` values before writing `reports/live-requests.jsonl` or `tmp-live/`; Gemini log paths are normalised to `models/{model}`; nothing under `docs/`, `examples/`, `tests/`, `sources/`, `generated/`, `reports/` may contain a key; raw responses live only in the gitignored `tmp-live/`; fixtures cut from live captures have `thoughtSignature`/`encrypted_content` truncated.5556## Checklist5758- [ ] Redaction layer before the provider call (keys, tokens, PII per your policy) and on tool outputs.59- [ ] `store: false` (OpenAI, Gemini Interactions) unless chaining/retrieval is needed; xAI ZDR toggled for regulated teams (with client-side state); ZDR-ineligible features inventoried per provider (Gemini grounding storage, xAI stateful features).60- [ ] Gemini: billing enabled (Paid Services) before any non-public data; free-tier keys only for public/synthetic prompts.61- [ ] Logs: headers masked (`Authorization`, `x-api-key`, `x-goog-api-key`), no `?key=` URLs, bodies not logged, request ids (`x-request-id`, `request-id`, `responseId`) kept; log retention short.62- [ ] Opaque hashed ids in `safety_identifier` (OpenAI/xAI) / `metadata.user_id` (Anthropic) / `labels.safety_identifier` (Gemini).63- [ ] Output filters for secret patterns and unexpected URLs (incl. xAI inline `[[N]](url)` citations).64- [ ] Files deleted after use (xAI: no auto-expiry unless `expires_after`; public URLs revoked); MCP/webhook third parties documented in your data map.65- [ ] Pre-commit secret scanning on the repo (patterns for all four key formats).66