SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.6 KB

# Computer use (computer_toolset_20260801, computer_20251124, computer_20250124)

Status: DOCUMENTED · LIVE_VERIFIED 2026-09-18 for request/response shapes (h1/h2 computer_20251124 on claude-sonnet-4-6; h3 computer_20250124 on haiku 4.5; h4 computer_toolset_20260801 on claude-sonnet-5; h5/k17/k17b/k22 error cases). No real display was driven — the loop below returns a 1×1 PNG placeholder. Sources: Computer use tool · Tool reference — client toolsets · Beta Messages reference · Release notes 2026-08-19. Last verified: 2026-09-18.

# Versions

type Beta header Request entry Call shape Models
computer_toolset_20260801 (GA 2026-08-19) none {"type":"computer_toolset_20260801", "configs"?: {member: {enabled?, defer_loading?}}, "cache_control"?, "allowed_callers"?: ["direct"]} — no name, no display_ fields* tool_use{name: <member>, toolset_name: "computer", input: {…}} Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 5, Sonnet 5, Opus 4.8 (live: haiku 400 "does not support tool types")
computer_20251124 computer-use-2025-11-24 {"type","name":"computer","display_width_px","display_height_px","display_number"?,"enable_zoom"?} tool_use{name:"computer", input:{action, …}} the 7 above + Opus 4.7, Opus 4.6, Sonnet 4.6, Opus 4.5 (live: sonnet 4.6 OK, haiku 400, no header → 400 unknown tag)
computer_20250124 computer-use-2025-01-24 same as above without enable_zoom same Sonnet 4.5, Haiku 4.5 (+ retired Opus 4.1 / Sonnet 4 / Opus 4) (live: haiku OK)
computer_20241022 computer-use-2024-10-22 legacy — Sonnet 3.5 (retired)

Platforms: toolset on Claude API + Google Cloud; AWS/Bedrock/Foundry only the beta versions. Not available in Managed Agents.

# Actions / members

Member (toolset) = input.action (2025 versions) Input Result you return
screenshot {} image block (PNG/JPEG base64)
zoom (toolset default on; 20251124 needs enable_zoom:true; absent in 20250124) region: [x0,y0,x1,y1] image (region at full res, scaled to screenshot size)
left_click, right_click, middle_click, double_click, triple_click coordinate?: [x,y], text?: modifiers `shift ctrl
left_click_drag start_coordinate, coordinate, text? OK
mouse_move coordinate OK
left_mouse_down, left_mouse_up {} OK
cursor_position {} text X=512, Y=384
scroll `scroll_direction: up down
type text OK
key text (e.g. Return, ctrl+s), repeat? 1–100 (toolset) OK
hold_key text, duration ≤ 300 s OK
wait duration ≤ 300 s OK

Coordinates are in the pixel space of the screenshot you returned (toolset) or of display_width_px × display_height_px (2025 versions). If you downscale screenshots, scale coordinates back up; macOS Retina = ×2. Image limits: Opus 4.7+ models 2576 px long edge / 4784 visual tokens; earlier 1568 px / ~1.15 MP; the toolset rejects oversized images instead of downscaling. > 20 images per request → stricter per-image limits.

Live shapes:

json
// h1 (computer_20251124, sonnet 4.6)
{"type":"tool_use","id":"toolu_0185…","name":"computer","input":{"action":"screenshot"},"caller":{"type":"direct"}}
// h4 (computer_toolset_20260801, sonnet 5) — preceded by an adaptive `thinking` block
{"type":"tool_use","id":"toolu_0181…","name":"screenshot","toolset_name":"computer","input":{},"caller":{"type":"direct"}}

# Batch actions (toolset) and results

Several member tool_use blocks per turn = batch. Run in order, stop at the first failure, answer every block: success → normal result; failed → is_error:true + description; later ones → is_error:true + exactly Not executed: an earlier computer action in this turn failed. Every member result must echo "toolset_name": "computer" and may contain only text/image blocks. cache_control on any tool_use/tool_result of a batch acts once at the end of the batch. Limit to one action per turn with tool_choice: {type: auto, disable_parallel_tool_use: true}.

Rejected on the toolset entry (400): name (k22: "name is not accepted on a toolset entry…"), display_*, enable_zoom, strict, input_examples, defer_loading on the entry, tool_choice type tool, legacy fine-grained header, a second entry / another tool named computer, configs disabling every member.

# Client-side loop (Python) — examples/anthropic/computer-use/computer_use_loop.py

python
def sampling_loop(prompt, max_iterations=10):
    messages = [{"role": "user", "content": prompt}]
    for _ in range(max_iterations):
        resp = client.beta.messages.create(model=MODEL, max_tokens=4096, tools=TOOLS, messages=messages,
                                           betas=["computer-use-2025-11-24"])      # no betas for computer_toolset_20260801
        messages.append({"role": "assistant", "content": resp.content})
        results, failed = [], False
        for b in resp.content:
            if b.type != "tool_use" or not (b.name == "computer" or getattr(b, "toolset_name", None) == "computer"):
                continue                                    # dispatch bash / text editor / custom tools here
            action = b.name if getattr(b, "toolset_name", None) else b.input["action"]
            r = {"type": "tool_result", "tool_use_id": b.id}
            if getattr(b, "toolset_name", None): r["toolset_name"] = "computer"
            if failed: r.update(content=NOT_EXECUTED, is_error=True)
            else:
                try: r["content"] = execute_action(action, b.input)   # screenshot -> [{"type":"image",...}], else "OK"
                except Exception as e: r.update(content=f"Error: {e}", is_error=True); failed = True
            results.append(r)
        if not results or resp.stop_reason != "tool_use": return messages
        messages.append({"role": "user", "content": results})

Ran live (2 requests each): step 1 → screenshot request; step 2 with a 1×1 PNG → Claude described a dark screen (h2, stopped at max_tokens 50). execute_action is a stub — connect it to Xvfb/xdotool in a sandboxed VM.

# Security, limits, pricing

  • Run in a dedicated VM/container, no credentials, domain allowlist, human confirmation for consequential actions; Anthropic's prompt-injection classifiers scan screenshots (opt-out via support). Screenshots/keystrokes never stored by Anthropic → ZDR-eligible (client tool).
  • Toolset definition ≈ 4,500 input tokens (h4: 4,189 with zoom disabled; sonnet 5); 2025 versions: 466–499 system-prompt tokens + ≈735 per definition (h1: 1,843 total; j8 count_tokens haiku: 1,826). Screenshots ≈ 1,000–1,800 tokens each. Prune screenshots in batches; on Fable 5.1 prefer server-side tool-result clearing (client pruning invalidates thinking blocks).
  • Thinking: models using the toolset omit thinking text by default — set thinking.display: "summarized" to inspect reasoning. For computer_20251124, docs suggest effort high (Opus 4.7) / medium (Sonnet 4.6, Opus 4.6).

Examples: computer_toolset_20260801.sh, computer_20251124.sh, computer_20250124_haiku.sh, computer_use_loop.py|.ts in examples/anthropic/computer-use/. Test: test_computer_use_beta_screenshot (expensive).