Provider Tool-Calling Verification — 2026-07-30
Phase 7.1 live sweep (ZyquoAgent --verify): every agent-capable model receives
the list_files tool schema, must emit a valid tool call (streamed), consume the
threaded tool_result, and produce a final answer naming the listed files.
80 models tested · 77 green · 3 failed (all 3 failures are provider-side — see the classification at the bottom)
This table merges the full pass-1 sweep with the per-provider re-runs performed after the two client fixes below (each fixed model was re-verified green live):
- OpenAI GPT-5.4/5.5/5.6:
/chat/completionsrejects function tools combined with anyreasoning_effortother than"none"— the client now downgradesreasoning_effortto"none"whenever tools are attached for these models (OpenAICompatibleClient.buildBody). - Gemini 3.x (compat endpoint): tool calls carry an
extra_content.google.thought_signaturethat MUST be echoed back verbatim when the call is threaded into history, or Gemini rejects the request (INVALID_ARGUMENT: Function call is missing a thought_signature).ToolCallnow carries the opaque signature; the client captures it from streaming fragments / non-streaming entries and re-attaches it on assistanttool_calls.
| Provider | Model | Tool call | Threading | Streaming | Latency | Notes |
|---|---|---|---|---|---|---|
| OpenAI | gpt-5.6-sol |
✅ | ✅ | ✅ | 3.0s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.6-terra |
✅ | ✅ | ✅ | 1.4s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.6-luna |
✅ | ✅ | ✅ | 1.6s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.5 |
✅ | ✅ | ✅ | 4.1s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.4 |
✅ | ✅ | ✅ | 1.8s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.4-mini |
✅ | ✅ | ✅ | 0.9s | pass 1 ❌ (tools+reasoning_effort rejected); green after client fix |
| OpenAI | gpt-5.2 |
✅ | ✅ | ✅ | 3.3s | |
| OpenAI | gpt-5.1 |
✅ | ✅ | ✅ | 1.6s | |
| OpenAI | gpt-5 |
✅ | ✅ | ✅ | 2.4s | |
| OpenAI | gpt-5-mini |
✅ | ✅ | ✅ | 2.1s | |
| OpenAI | o3 |
✅ | ✅ | ✅ | 1.3s | |
| OpenAI | o4-mini |
✅ | ✅ | ✅ | 1.2s | |
| Anthropic | claude-opus-5 |
✅ | ✅ | ✅ | 3.9s | |
| Anthropic | claude-sonnet-5 |
✅ | ✅ | ✅ | 4.0s | |
| Anthropic | claude-fable-5 |
✅ | ✅ | ✅ | 7.5s | |
| Anthropic | claude-opus-4-8 |
✅ | ✅ | ✅ | 3.2s | |
| Anthropic | claude-opus-4-7 |
✅ | ✅ | ✅ | 3.2s | |
| Anthropic | claude-opus-4-6 |
✅ | ✅ | ✅ | 6.1s | |
| Anthropic | claude-sonnet-4-6 |
✅ | ✅ | ✅ | 4.1s | |
| Anthropic | claude-haiku-4-5-20251001 |
✅ | ✅ | ✅ | 1.6s | |
| xAI | grok-4.5 |
✅ | ✅ | ✅ | 2.5s | |
| xAI | grok-4.3 |
✅ | ✅ | ✅ | 4.5s | |
| xAI | grok-4.20 |
✅ | ✅ | ✅ | 4.9s | |
| xAI | grok-4.20-non-reasoning |
✅ | ✅ | ✅ | 1.2s | |
| xAI | grok-code-fast-1 |
✅ | ✅ | ✅ | 5.3s | |
| Mistral | mistral-medium-latest |
✅ | ✅ | ✅ | 1.4s | |
| Mistral | mistral-large-latest |
✅ | ✅ | ✅ | 3.7s | |
| Mistral | mistral-small-latest |
✅ | ✅ | ✅ | 0.9s | |
| Google Gemini | gemini-3.6-flash |
✅ | ✅ | ✅ | 2.1s | pass 1 ❌ (thought_signature missing on threading); green after client fix |
| Google Gemini | gemini-3.5-flash |
✅ | ✅ | ✅ | 2.0s | pass 1 ❌ (thought_signature); green after client fix |
| Google Gemini | gemini-3.5-flash-lite |
✅ | ✅ | ✅ | 1.0s | pass 1 ❌ (thought_signature); green after client fix |
| Google Gemini | gemini-3.1-pro-preview |
✅ | ✅ | ✅ | 3.9s | pass 1 ❌ (thought_signature); green after client fix |
| Google Gemini | gemini-2.5-pro |
✅ | ✅ | ✅ | 3.3s | |
| Google Gemini | gemini-2.5-flash |
✅ | ✅ | ✅ | 1.3s | |
| Google Gemini | gemini-pro-latest |
✅ | ✅ | ✅ | 4.2s | alias of Gemini 3.x; pass 1 ❌ (thought_signature); green after client fix |
| Google Gemini | gemini-flash-latest |
✅ | ✅ | ✅ | 1.5s | alias of Gemini 3.x; pass 1 ❌ (thought_signature); green after client fix |
| Alibaba Qwen | qwen3.7-max |
✅ | ✅ | ✅ | 3.0s | |
| Alibaba Qwen | qwen3.7-plus |
✅ | ✅ | ✅ | 2.6s | |
| Alibaba Qwen | qwen3.7-flash |
✅ | ✅ | ✅ | 1.6s | |
| Alibaba Qwen | qwen3.6-plus |
✅ | ✅ | ✅ | 2.2s | |
| Alibaba Qwen | qwen3.5-plus |
✅ | ✅ | ✅ | 2.0s | |
| Alibaba Qwen | qwen3-coder-plus |
✅ | ✅ | ✅ | 2.1s | |
| Alibaba Qwen | qwen3-coder-flash |
✅ | ✅ | ✅ | 2.0s | |
| Alibaba Qwen | qwen3-coder-next |
✅ | ✅ | ✅ | 2.3s | |
| Alibaba Qwen | qwen3-coder-480b-a35b-instruct |
✅ | ✅ | ✅ | 2.6s | |
| Alibaba Qwen | qwen3.5-397b-a17b |
✅ | ✅ | ✅ | 1.9s | |
| Alibaba Qwen | deepseek-v4-pro |
✅ | ✅ | ✅ | 4.6s | |
| Alibaba Qwen | glm-5.2 |
✅ | ✅ | ✅ | 2.1s | |
| DeepSeek | deepseek-v4-flash |
✅ | ✅ | ✅ | 2.5s | |
| DeepSeek | deepseek-v4-pro |
✅ | ✅ | ✅ | 3.2s | |
| Kimi | kimi-k3 |
❌ | ❌ | ❌ | 100.5s | INFRA: HTTP 429 engine_overloaded_error on most requests during the sweep window (reproduced with raw curl in every request shape; intermittent 200s prove the key/tier is fine). One harness attempt DID emit a correct list_files call, so the wire format is right — Moonshot capacity issue; re-verify when the engine recovers |
| Kimi | kimi-k2.7-code |
✅ | ✅ | ✅ | 2.3s | |
| Kimi | kimi-k2.7-code-highspeed |
✅ | ✅ | ✅ | 1.8s | |
| Kimi | kimi-k2.6 |
✅ | ✅ | ✅ | 5.1s | |
| Kimi | kimi-k2.5 |
✅ | ✅ | ✅ | 4.8s | |
| Together AI | moonshotai/Kimi-K3 |
✅ | ✅ | ✅ | 19.0s | slow first token (reasoning) |
| Together AI | moonshotai/Kimi-K2.7-Code |
✅ | ✅ | ✅ | 1.6s | |
| Together AI | deepseek-ai/DeepSeek-V4-Pro |
✅ | ✅ | ✅ | 1.6s | |
| Together AI | zai-org/GLM-5.2 |
✅ | ✅ | ✅ | 1.2s | |
| Together AI | Qwen/Qwen3.7-Max |
✅ | ✅ | ✅ | 4.6s | |
| Together AI | openai/gpt-oss-120b |
✅ | ✅ | ✅ | 2.2s | |
| Together AI | nvidia/nemotron-3-ultra-550b-a55b |
✅ | ✅ | ✅ | 1.2s | |
| Together AI | MiniMaxAI/MiniMax-M3 |
✅ | ✅ | ✅ | 6.1s | |
| DeepInfra | anthropic/claude-fable-5 |
✅ | ✅ | ✅ | 6.3s | |
| DeepInfra | anthropic/claude-opus-5 |
✅ | ✅ | ✅ | 5.6s | |
| DeepInfra | anthropic/claude-sonnet-5 |
✅ | ✅ | ✅ | 7.1s | |
| DeepInfra | anthropic/claude-opus-4-8 |
✅ | ✅ | ✅ | 4.1s | |
| DeepInfra | anthropic/claude-haiku-4-5 |
✅ | ✅ | ✅ | 2.3s | |
| DeepInfra | google/gemini-3.1-pro |
✅ | ❌ | ❌ | 4.0s | SERVER-SIDE: step 1 works, but DeepInfra never returns Gemini's thought_signature and strips extra_content from requests (verified with raw probes, incl. Google's documented bypass placeholder) → Google rejects every threaded tool result with INVALID_ARGUMENT. Multi-turn tool use is impossible through DeepInfra — REMOVAL CANDIDATE (the same model is green via the native Gemini provider) |
| DeepInfra | google/gemini-3.5-flash |
✅ | ❌ | ❌ | 2.3s | SERVER-SIDE: same as google/gemini-3.1-pro — thought_signature dropped by DeepInfra; REMOVAL CANDIDATE |
| DeepInfra | deepseek-ai/DeepSeek-V4-Pro |
✅ | ✅ | ✅ | 4.1s | |
| DeepInfra | deepseek-ai/DeepSeek-V4-Flash |
✅ | ✅ | ✅ | 3.3s | |
| DeepInfra | moonshotai/Kimi-K2.7-Code |
✅ | ✅ | ✅ | 2.2s | |
| DeepInfra | zai-org/GLM-5.2 |
✅ | ✅ | ✅ | 3.1s | |
| DeepInfra | Qwen/Qwen3.7-Max |
✅ | ✅ | ✅ | 4.7s | |
| DeepInfra | Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo |
✅ | ✅ | ✅ | 0.8s | |
| DeepInfra | openai/gpt-oss-120b |
✅ | ✅ | ✅ | 2.6s | |
| DeepInfra | MiniMaxAI/MiniMax-M3 |
✅ | ✅ | ✅ | 3.0s | |
| Cerebras | gpt-oss-120b |
✅ | ✅ | ✅ | 1.4s | |
| Cerebras | gemma-4-31b |
✅ | ✅ | ✅ | 0.6s |
Failure classification
| Model | Class | Evidence / disposition |
|---|---|---|
deepinfra google/gemini-3.1-pro |
Broken server-side | DeepInfra never emits thought_signature and drops extra_content on requests; Gemini 3 requires it for threaded tool calls. No client-side workaround exists (Google's documented placeholder is also stripped). Recommend removing from the agent-capable set — the identical models are green via the native Gemini provider. |
deepinfra google/gemini-3.5-flash |
Broken server-side | Same as above. Recommend removing from the agent-capable set. |
kimi kimi-k3 |
Infra (provider capacity) | HTTP 429 engine_overloaded_error on most requests during the sweep window, reproduced with raw curl in every request shape; intermittent 200s prove the key/tier is fine and one harness attempt did emit a correct list_files call. Not a wire-format or capability failure — keep in the set, re-verify when Moonshot capacity recovers. |
No model id in the agent-capable set was rejected as nonexistent/deprecated, and no model failed tool calling because of model capability — every ❌ above is provider-side.