SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.6 KB

# Computer use (GUI agents)

Status: DOCUMENTED (Gemini computerUse on generateContent ACCOUNT_RESTRICTED on this key — 429 limit: 0; the Interactions {"type":"computer_use"} declaration was accepted live 2026-09-18; xAI has no computer-use tool) Sources: OpenAI OpenAPI spec (ComputerToolCall.pending_safety_checks, acknowledged_safety_checks) · https://developers.openai.com/api/docs/guides/tools-computer-use + …-integration · https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool · …/browser-use-tool · Gemini: https://ai.google.dev/gemini-api/docs/computer-use (tools:[{computerUse:{environment: ENVIRONMENT_BROWSER, excludedPredefinedFunctions[], disabledSafetyPolicies[], enablePromptInjectionDetection}}]; predefined actions open_web_browser, navigate, click_at, type_text_at, key_combination, scroll_document, drag_and_drop…; functionCall.args.safety_decision = {decision: "require_confirmation" | "blocked", explanation} → reply safety_acknowledgement: true only after the user confirms; reference sandbox github.com/google/computer-use-preview; Preview "may contain errors and security vulnerabilities") · https://ai.google.dev/api/generate-content (#ComputerUse #SafetyPolicy: FINANCIAL_TRANSACTIONS, SENSITIVE_DATA_MODIFICATION, COMMUNICATION_TOOL, ACCOUNT_CREATION, DATA_MODIFICATION, USER_CONSENT_MANAGEMENT, LEGAL_TERMS_AND_AGREEMENTS) · xAI: https://docs.x.ai/developers/tools/overview (no computer-use tool; shell local tool only) Last verified: 2026-09-19

# Why it is the highest-risk tool

The model sees screenshots/page text (untrusted, injectable) and emits clicks/keystrokes/commands that you execute on a real machine. Anthropic: "instructions on webpages or contained in images might override your instructions or cause Claude to make mistakes." Both vendors' guidance converges on the same four controls.

# Provider mechanics

# OpenAI (computer_use_preview / integration guide)

  • The model returns computer_call items. Each may carry pending_safety_checks[] ({id, code, message}), with codes such as malicious_instructions (suspected injection in the screen content), irrelevant_domain (navigating off-task), sensitive_domain (e.g. banking). The docs require you to surface these to a human and, only if confirmed, echo them back as acknowledged_safety_checks in the computer_call_output — otherwise the model will not proceed.
  • Integration examples run actions through an execute_in_sandbox boundary that "must preserve the browser or desktop session, enforce execution limits, and apply your permission rules".
  • Confirmation guidance: confirm at the point of risk (right before the purchase/send/delete), choose the confirmation level by consequence, and hand off to a human for logins, CAPTCHAs and payments.

# Anthropic (computer_toolset_20260801, browser_toolset_20260801, legacy computer_20250124)

  • Documented precautions (verbatim intent): (1) dedicated VM/container with minimal privileges; (2) no access to sensitive data such as account logins; (3) internet limited to an allowlist of domains; (4) human confirmation for consequential actions and anything needing affirmative consent (cookies, payments, ToS).
  • Classifiers automatically scan tool returns (screenshots, page text) for prompt injection and can steer Claude to ask the user; opt-out only via support and "won't be ideal for use cases without a human in the loop".
  • Batched actions: every tool_use block must be answered; on a failed action return is_error: true and halt the rest of the batch so the model re-plans with true state.
  • If a login is unavoidable, credentials go in the prompt inside tags like <robot_credentials> — the docs flag that this increases risk; prefer pre-authenticated sessions with the minimum scope.
  • Reference implementation runs inside Docker with port mappings for observation.

# Gemini (computerUse, Preview — gemini-2.5-computer-use-preview-10-2025 on generateContent, Gemini 3 via Interactions)

  • The model emits ordinary functionCall parts for predefined browser actions (click_at {x, y} on a 0–999 grid, type_text_at, navigate {url}, key_combination, scroll_document, drag_and_drop…); you execute them (Playwright etc.) and answer with a functionResponse whose response carries the new url and a screenshot (parts[].inlineData image). Echo the whole model turn (thoughtSignature, id).
  • Safety decisions: a call may carry args.safety_decision = {"decision": "require_confirmation" | "blocked", "explanation": "…"}. The docs require you to prompt the user and, only if they confirm, include "safety_acknowledgement": true in the functionResponse.response; on blocked (or a refusal) terminate the task. Implement this regardless of disabledSafetyPolicies — those are preferences (the model may still ask), not a way to switch checks off.
  • excludedPredefinedFunctions removes actions you never want (e.g. drag_and_drop, key_combination); enablePromptInjectionDetection turns on Google's injection detector for screen content; custom safety instructions in systemInstruction are recommended; the tool is not available in the Live API.
  • Run inside the reference Docker sandbox or an equivalent VM; pricing is regular token pricing (screenshots = image tokens) and not available on the free tier.

# xAI

No computer-use tool exists. The closest primitives are the local shell tool (code-execution-and-sandboxing.md) and X/web browsing on xAI's side (x_search, web_search — not a GUI agent). If you build a GUI agent on Grok with function calling, all controls below are yours to implement; there is no pending_safety_checks/safety_decision equivalent.

# Your controls (all providers)

Layer Control
Isolation ephemeral VM/container per session; no host mounts; snapshot & destroy
Network egress allowlist (DNS + HTTP proxy) — the model must not reach internal services, cloud metadata (169.254.169.254), or arbitrary exfil hosts
Identity throwaway accounts; never the operator's session; no password manager, no saved cards
Actions deny-list dangerous key combos / URLs client-side; require confirmation for payments, sends, deletes, downloads-then-execute
Observation record screenshots + actions + safety-check ids per step (audit trail)
Budget max steps, max wall time, max_uses-style caps; kill switch
Injection treat every screenshot as attacker-controlled; keep OpenAI safety checks and Anthropic classifiers on

# Checklist

  • Dedicated sandbox VM/container, destroyed after the task.
  • Domain allowlist enforced at the network layer, not just in the prompt.
  • No real credentials/cards in the environment; pre-authenticated minimal-scope sessions if needed.
  • OpenAI: pending_safety_checks shown to a human; acknowledged_safety_checks only after explicit confirmation; never auto-acknowledge.
  • Anthropic: classifier layer left on; is_error + halt on failed batch actions.
  • Gemini: safety_decision handled on every call (require_confirmation → human; blocked → stop); safety_acknowledgement: true never sent automatically; enablePromptInjectionDetection on; excludedPredefinedFunctions pruned; disabledSafetyPolicies left empty unless reviewed.
  • xAI / custom GUI agents without provider safety checks: your own confirmation layer before consequential actions.
  • Human confirmation at the point of risk for consequential actions.
  • Step/time/cost limits and a kill switch; full action log retained.