SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
8.9 KB

# Gemini — Computer Use (tools[].computerUse, Preview)

Status: DOCUMENTED + ACCOUNT_RESTRICTED (2026-09-18: our key got HTTP 429 quota errors — see "Live verification" section at the end)

Sources:

Last verified: 2026-09-18 (docs only)

# 1. Models

Model Notes
gemini-3.8-flash recommended; high-accuracy UI interaction
gemini-3.7-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3-flash-preview Gemini 3.x feature set (intents, 3 environments, safety policies, prompt-injection detection from 3.5 Flash)
gemini-3.6-flash, gemini-robotics-er-2-preview, gemini-robotics-er-1.6-preview "Supported (Preview)" on model pages, not in the guide list
gemini-2.5-computer-use-preview-10-2025 legacy; browser only; 128k input / 64k output; image + text input
gemini-3-pro-preview model page: Not supported (changelog 2025-11 said launched) — conflict

# 2. Tool config (tools[].computerUse)

Field Type Notes
environment enum, required ENVIRONMENT_BROWSER, ENVIRONMENT_MOBILE, ENVIRONMENT_DESKTOP (ENVIRONMENT_UNSPECIFIED → browser). Mobile/desktop: Gemini 3.x
excludedPredefinedFunctions[] string[] drop predefined actions (restrict action space or replace with your own declaration)
enablePromptInjectionDetection bool (default false) Gemini 3.5 Flash+: scan screenshots for hidden instructions; blocks execution when detected
disabledSafetyPolicies[] SafetyPolicy[] FINANCIAL_TRANSACTIONS, SENSITIVE_DATA_MODIFICATION, COMMUNICATION_TOOL, ACCOUNT_CREATION, DATA_MODIFICATION, USER_CONSENT_MANAGEMENT, LEGAL_TERMS_AND_AGREEMENTS — preferences only; require_confirmation may still be returned

Extra: add tools[].functionDeclarations for custom actions (e.g. yield_to_user(reason)); tune generationConfig.thinkingConfig.thinkingLevel (lower = faster). No display size needed — coordinates are normalized 0–999 and you scale them to your viewport.

# 3. Predefined actions (functionCall names)

# Browser (ENVIRONMENT_BROWSER, Gemini 3.x)

Action Args (all + intent: str)
click, double_click, triple_click, middle_click, right_click, mouse_down, mouse_up, move x, y (0–999)
type text, press_enter (default false)
drag_and_drop start_x, start_y, end_x, end_y
wait seconds (default 1)
press_key, key_down, key_up key
hotkey keys: list[str]
take_screenshot —
scroll x, y, direction: up|down|left|right, magnitude_in_pixels (default 300)
go_back, go_forward —
navigate url

# Mobile (ENVIRONMENT_MOBILE): open_app(app_name), click, list_apps, wait, go_back, type, drag_and_drop, long_press(x,y,seconds=2), press_key, take_screenshot.

# Desktop (ENVIRONMENT_DESKTOP): browser set minus go_back/navigate/go_forward.

# Legacy (gemini-2.5-computer-use-preview-10-2025)

Action Args
open_web_browser, wait_5_seconds, go_back, go_forward, search none
navigate url
click_at, hover_at x, y (0–999)
type_text_at x, y, text, press_enter (default true), clear_before_typing (default true)
key_combination keys e.g. "Control+A"
scroll_document direction
scroll_at x, y, direction, magnitude (default 800)
drag_and_drop x, y, destination_x, destination_y

# 4. Safety decisions

Model: functionCall.args.safety_decision = {"decision": "require_confirmation", "explanation": "..."} (regular/allowed, require_confirmation, or blocked). Client: prompt the user; if confirmed, include "safety_acknowledgement": true in the functionResponse response object; otherwise terminate. Always implement this handling regardless of disabledSafetyPolicies.

# 5. Loop (pseudo-code, Playwright)

python
from playwright.sync_api import sync_playwright
from google import genai; from google.genai import types
client = genai.Client(); W, H = 1440, 900
cfg = types.GenerateContentConfig(tools=[types.Tool(computer_use=types.ComputerUse(environment="ENVIRONMENT_BROWSER"))])
with sync_playwright() as p:
    page = p.chromium.launch().new_context(viewport={"width": W, "height": H}).new_page(); page.goto("https://www.google.com")
    history = [types.Content(role="user", parts=[types.Part(text="Search for 'Gemini API'."),
               types.Part.from_bytes(data=page.screenshot(type="png"), mime_type="image/png")])]
    for _ in range(10):
        r = client.models.generate_content(model="gemini-3.8-flash", contents=history, config=cfg)
        history.append(r.candidates[0].content)                     # keeps thoughtSignature + id
        calls = [pt.function_call for pt in r.candidates[0].content.parts if pt.function_call]
        if not calls: break
        responses = []
        for fc in calls:
            a = dict(fc.args); result = {}
            if "safety_decision" in a: result["safety_acknowledgement"] = True   # after asking the user!
            if fc.name in ("click", "click_at"): page.mouse.click(a["x"] / 1000 * W, a["y"] / 1000 * H)
            elif fc.name in ("type", "type_text_at"):
                if "x" in a: page.mouse.click(a["x"] / 1000 * W, a["y"] / 1000 * H)
                page.keyboard.type(a["text"]);  a.get("press_enter") and page.keyboard.press("Enter")
            elif fc.name == "navigate": page.goto(a["url"])
            elif fc.name == "scroll": page.mouse.wheel(0, a.get("magnitude_in_pixels", 300) * (1 if a["direction"] == "down" else -1))
            page.wait_for_load_state("load", timeout=5000)
            responses.append(types.Part(function_response=types.FunctionResponse(name=fc.name, id=fc.id,
                response={"url": page.url, **result},
                parts=[types.FunctionResponsePart(inline_data=types.FunctionResponseBlob(mime_type="image/png", data=page.screenshot(type="png")))])))
        history.append(types.Content(role="user", parts=responses))

The official guide's get_function_responses sample returns the screenshot as an image content block alongside {"url": current_url, ...}; the generateContent shape is functionResponse.parts[].inlineData (Gemini 3 multimodal function responses).

# 6. Pricing / limits / security

  • Pricing: regular model token pricing (screenshots = image tokens); free tier not available; legacy 2.5 CU model has its own table.
  • Preview: "may contain errors and security vulnerabilities" — avoid critical decisions, sensitive data, irreversible actions.
  • Run in a sandboxed VM/container (reference Docker sandbox); human-in-the-loop on require_confirmation; custom safety system instructions; enable prompt-injection detection where available; not in the Live API.

# 7. Minimal request

bash
curl -s "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
  -d '{"contents":[{"parts":[{"text":"Search for Gemini API on Google."}]}],"tools":[{"computerUse":{"environment":"ENVIRONMENT_BROWSER"}}]}'
ts
const r = await ai.models.generateContent({ model: 'gemini-3.5-flash-lite', contents: "Search for 'Gemini API' on Google.",
  config: { tools: [{ computerUse: { environment: 'ENVIRONMENT_BROWSER', excludedPredefinedFunctions: ['drag_and_drop'] } }] } });
console.log(r.functionCalls);

# Live verification (2026-09-18)

  • Legacy generateContent tool tools:[{"computerUse":{"environment":"ENVIRONMENT_BROWSER"}}] on gemini-2.5-computer-use-preview-10-2025 with a 1×1 PNG → HTTP 429 Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_input_token_count, limit: 0, model: computer-use-preview → ACCOUNT_RESTRICTED (tmp-live/gemini-tools/f_computer_use_generateContent.json).
  • Interactions API tools:[{"type":"computer_use","environment":"browser"}] on gemini-3.5-flash-lite, input "Do nothing. Reply DONE." → HTTP 200, status:"completed", steps [thought, model_output "DONE"] — the tool declaration is accepted on a non-computer-use model; no action was requested so the function_call shape was not observed live (f2_computer_use_interactions.json).
  • The client loop (Playwright, function_result with screenshot + url, safety_decision acknowledgement) in examples/gemini/tools/computer-use/computer_use_interactions.py --loop is DOCUMENTED, not executed.