SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
10.0 KB

# Gemini tool loop — generateContent, Interactions API and Live API

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-18, gemini-3.5-flash-lite for REST, gemini-2.5-flash-native-audio-latest for Live; runnable loops examples/shared/tool-loop/gemini_tool_loop.{py,ts}, examples/gemini/interactions/interactions_tools.py, raw responses tmp-live/gemini-tools/). Sources: Function calling · Thought signatures · Interactions API reference · Live API tools · Tool combination · Live reference. Last verified: 2026-09-18. Sibling pages: OpenAI tool loop, Anthropic tool loop; tool catalogue docs/tools/gemini/index.md.

# 1. Three surfaces, one loop

Surface Where state lives Call marker Your reply Stop condition
POST /v1beta/models/{model}:generateContent (+ :streamGenerateContent) client — you resend the whole contents[] candidates[0].content.parts[] containing functionCall{name,args,id} (+ thoughtSignature on the part) user Content whose parts are functionResponse{name, response, id?} — one part per call, all in one Content a model turn with no functionCall part (finishReason: STOP)
POST /v1beta/interactions server — chain with previous_interaction_id (or resend input[] as Steps) status: "requires_action" and steps[] items {type:"function_call", id:"call_…", name, arguments} new interaction whose input is [{type:"function_result", name, call_id, result:[{type:"text", text}]}] status: "completed" (or failed/cancelled)
Live API (BidiGenerateContent WebSocket) session (server) server message toolCall{functionCalls:[{name,args,id:"function-call-…"}]} client message toolResponse{functionResponses:[{name, response, id, scheduling?, willContinue?}]} model keeps generating; toolCallCancellation{ids} if the user interrupts
sequenceDiagram
  participant App
  participant Gemini
  App->>Gemini: contents[user] + tools[functionDeclarations] (+ toolConfig.functionCallingConfig)
  Gemini-->>App: model turn: parts[functionCall+thoughtSignature, functionCall…]
  loop every functionCall of the turn
    App->>App: run handler(args) → JSON result (or {"error": …})
  end
  App->>Gemini: contents[user, model turn VERBATIM, user{functionResponse ×N}]
  Gemini-->>App: text (or another functionCall turn)

# 2. Rules confirmed live (2026-09-18)

  1. Echo the model turn verbatim. Each functionCall part comes with a sibling thoughtSignature. Dropping it → 400 Function call is missing a thought_signature in functionCall parts … function call default_api:get_weather, position 2 (tmp-live/gemini-tools/a2b_fc_no_thought_signature.json). In a parallel turn only the first call carries a signature — copy the parts array as-is and you are compliant.
  2. Answer every call of the turn, in one user Content. Parallel turn [get_weather, get_time] → one {"role":"user","parts":[{functionResponse…},{functionResponse…}]}. functionResponse.response must be a JSON object (wrap scalars: {"result": …}); the id returned in functionCall (call_…) may be echoed in functionResponse.id.
  3. Forcing: toolConfig.functionCallingConfig.mode = AUTO (default) | ANY (must call; optionally allowedFunctionNames) | NONE | VALIDATED (schema-validated; accepted live on 3.5-flash-lite). Reset to AUTO after the first forced turn or the model will call forever.
  4. Schemas: parameters uses the OpenAPI subset with upper-case types (OBJECT, STRING…); parametersJsonSchema accepts plain JSON Schema (enum, lower-case types) — both verified. response/responseJsonSchema describe the tool output (optional).
  5. behavior: NON_BLOCKING is Live-only: generateContent → 400 FunctionDeclaration.behavior is only supported by the BidiGenerateContent method. In Live the model kept speaking while the toolCall was outstanding; toolResponse.functionResponses[].scheduling (INTERRUPT | WHEN_IDLE | SILENT) decides how the late result is folded in. setup.toolConfig is rejected in Live (close 1007) — no forcing there.
  6. Built-in tools mixed with functions: googleSearch/codeExecution/urlContext/fileSearch/googleMaps run server-side within the same turn (their parts — executableCode, codeExecutionResult, groundingMetadata, urlContextMetadata — need no reply); Gemini 3 combines them with functionDeclarations (toolConfig.includeServerSideToolInvocations circulates the server tool context). See tool-combination.
  7. Interactions API differences: snake_case everywhere (generation_config, previous_interaction_id; camelCase → 400 "Did you mean 'generation_config'?"); tool choice lives in generation_config.tool_choice ({"allowed_tools":{"mode":"any","tools":[…]}} or the bare enum string) — top-level tool_choice → 400; results are Steps (function_result{name, call_id, result:[Content]}); store:false disables chaining (no id).
  8. Budget the loop: bound turns (6 in the examples), pass tool errors back as data, keep maxOutputTokens small while iterating, and count usageMetadata.toolUsePromptTokenCount (server-tool retrieved content billed as input).

# 3. Minimal loop (Python, REST)

python
from scripts.live import gemini_request
MODEL = "gemini-3.5-flash-lite"
def run(prompt, declarations, handlers, max_turns=6):
    contents = [{"role": "user", "parts": [{"text": prompt}]}]
    for _ in range(max_turns):
        st, r, _ = gemini_request("POST", f"/v1beta/models/{MODEL}:generateContent",
                                  {"contents": contents, "tools": [{"functionDeclarations": declarations}], "generationConfig": {"maxOutputTokens": 256}})
        turn = r["candidates"][0]["content"]                       # keep thoughtSignature parts intact
        calls = [p["functionCall"] for p in turn["parts"] if "functionCall" in p]
        if not calls:
            return "".join(p.get("text", "") for p in turn["parts"])
        replies = [{"functionResponse": {"name": c["name"], "id": c.get("id"), "response": handlers[c["name"]](c.get("args", {}))}} for c in calls]
        contents += [turn, {"role": "user", "parts": replies}]

Full versions with logging: examples/shared/tool-loop/gemini_tool_loop.py / .ts (executed 2026-09-18: get_weather + get_time in one parallel turn, then the final sentence).

# 4. Interactions API loop (server-side state)

python
body = {"model": MODEL, "input": prompt, "tools": [{"type": "function", "name": ..., "parameters": {...}}]}
while True:
    st, r, _ = gemini_request("POST", "/v1beta/interactions", body)
    calls = [s for s in r["steps"] if s["type"] == "function_call"]
    if r["status"] != "requires_action" or not calls:
        break                                                      # r["steps"] holds model_output content
    body = {"model": MODEL, "previous_interaction_id": r["id"], "tools": body["tools"],
            "input": [{"type": "function_result", "name": c["name"], "call_id": c["id"],
                       "result": [{"type": "text", "text": json.dumps(handlers[c["name"]](c["arguments"]))}]} for c in calls]}

Verified in examples/gemini/interactions/interactions_tools.py (requires_action → completed). Deep Research / Antigravity agents run the loop server-side (background:true, poll GET /v1beta/interactions/{id} or stream).

# 5. Live API loop

Client sends setup{tools:[{functionDeclarations:[…]}]}; on toolCall run the handlers and send {"toolResponse":{"functionResponses":[{"name":…,"id":…,"response":{…}}]}}. With behavior: NON_BLOCKING add "scheduling":"WHEN_IDLE" (or INTERRUPT/SILENT) so the model decides when to voice the result; willContinue:true announces further results for the same id. Handle toolCallCancellation{ids} (user barge-in) by discarding in-flight work. Observed message order: live-events.md.

# 6. Cross-provider mapping

Concept Gemini generateContent Gemini Interactions OpenAI Responses Anthropic Messages
declaration tools[].functionDeclarations[]{name,description,parameters} tools[]{type:"function",name,parameters} tools[]{type:"function",name,parameters} tools[]{name,input_schema}
force toolConfig.functionCallingConfig.mode:"ANY" generation_config.tool_choice:"any" tool_choice:"required" tool_choice:{type:"any"}
call item part functionCall{name,args,id} (+thoughtSignature) step function_call{id,name,arguments} item function_call{call_id,name,arguments} block tool_use{id,name,input}
result item part functionResponse{name,response,id} input function_result{name,call_id,result[]} function_call_output{call_id,output} tool_result{tool_use_id,content}
state client resends contents previous_interaction_id previous_response_id / conversation client resends messages
reasoning carry-over thoughtSignature (mandatory) thought step signature reasoning items / encrypted content thinking blocks (signature)

# Live verification (2026-09-18)

examples/shared/tool-loop/gemini_tool_loop.py and .ts executed (one parallel round trip each, HTTP 200 ×2); rules 1–5 and 7 come from the probes listed in docs/tools/gemini/function-calling.md and interactions-api.md. Not exercised: Live toolResponse round trip (the toolCall was observed, the reply was not sent), scheduling/willContinue, toolCallCancellation.