# Let Claude Draw on a Daytona Sandbox Desktop

This guide walks through the [`claude-draws`](https://github.com/daytona/guides/tree/main/python/computer-use/claude-draws) example: a Daytona sandbox boots a desktop, a self-contained HTML sketchpad opens in a visible Chromium window, and Claude is handed the keyboard and mouse through the Anthropic SDK's computer toolset (`computer_toolset_20260801`).

Nothing about the app is special-cased for the model. There is no DOM access, no accessibility tree, no scripted click targets. Claude takes a screenshot, finds the palette, scrolls it sideways, clicks a swatch by its visible label, and drags to draw — the same moves a person would make. The bridge between the model's tool calls and the sandbox is `DaytonaComputer` from the `daytona-claude-toolsets` package, which turns each member of the toolset into a call against Daytona's [Computer Use API](https://www.daytona.io/docs/en/computer-use.md).

---

### 1. Workflow Overview

```text
Your process                              Daytona sandbox
──────────────                            ───────────────
DaytonaComputer(...)         ──────────▶  create sandbox, start desktop (1280x800)
upload paint.html            ──────────▶  /tmp/claude-draws/paint.html
launch Chromium (--app=...)  ──────────▶  visible window, no tab strip or omnibox
                                   │
client.beta.messages.tool_runner   │
  model ⇄ computer toolset         │
    screenshot  ───────────────────┼───▶  screenshot API
    scroll / click / drag  ────────┼───▶  mouse API
    key / type  ───────────────────┼───▶  keyboard API
                                   ▼
close()                      ──────────▶  sandbox deleted
```

The model loop, your API keys, and the toolset all stay in your process. The sandbox only ever sees input events and hands back PNGs.

:::note[Throwaway by design]
In this example the desktop is a disposable sandbox created for the run and deleted when the `with` block exits, which is what makes it reasonable to approve every action the model asks for without a human at the keyboard. Point the same loop at a sandbox you intend to keep and that reasoning no longer holds.
:::

### 2. Project Setup

#### Requirements

- Python 3.10 or higher on your machine.
- A Daytona account and API key.
- An Anthropic API key.

#### Clone the Repository

Clone the [Daytona guides repository](https://github.com/daytona/guides) and set up the example:

```bash
git clone https://github.com/daytona/guides.git
cd guides/python/computer-use/claude-draws
python3 -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
pip install -e .
```

`pip install -e .` pulls in `daytona-claude-toolsets` (which supplies `DaytonaComputer`) and an `anthropic` release new enough to ship `anthropic.tools.computer`. `daytona-claude-toolsets` in turn requires the `daytona` SDK 0.223.x — the first release whose Python client exposes the mouse and keyboard hold calls the driver routes through.

#### Configure API Keys

```bash
cp .env.example .env
```

Fill in:

- `DAYTONA_API_KEY`: required. Get it from the [Daytona Dashboard](https://app.daytona.io/dashboard/keys).
- `ANTHROPIC_API_KEY`: required, for the model loop.
- `MODEL`: optional, the model the drawing loop uses. Defaults to `claude-sonnet-5-5`.

:::note[`.env` is a POSIX shell file]
The template is a list of `export` lines and nothing loads it for you, so `source .env` before running. PowerShell cannot source it — set the same variables in the session instead. Leave `MODEL` unset to take the default:

```powershell
$env:DAYTONA_API_KEY = "..."
$env:ANTHROPIC_API_KEY = "..."
# $env:MODEL = "..."  # only to pick a different model
```
:::

### 3. Opening the Sketchpad

`DaytonaComputer` is a context manager. Entering it creates the sandbox, starts its desktop, waits for a display to report a size, and exposes the live `Sandbox` object so you can set the screen up before the model ever looks at it:

```python
with DaytonaComputer(confirm=confirm, resolution=(WIDTH, HEIGHT)) as computer:
    print(f"sandbox {computer.sandbox.id}, screen {computer.width}x{computer.height}")
    upload_sketchpad(computer.sandbox)
    launch_chromium(computer.sandbox)
```

`paint.html` is uploaded with the sandbox filesystem API, then opened in Chromium in app mode (`--app=file://...`), so the window is nothing but the sketchpad — no tab strip, no omnibox, no first-run bubbles to confuse the model.

One detail is worth copying into your own runs. Chromium is a foreground GUI process that never returns, so launching it with `process.exec` would block until the browser exits. The example starts it through an async process session instead, then polls Chromium's loopback debugging port purely as a readiness probe:

```python
sandbox.process.create_session(SESSION_ID)
sandbox.process.execute_session_command(
    SESSION_ID, SessionExecuteRequest(command=command, run_async=True)
)
deadline = time.monotonic() + 60
probe = f"curl -sf --max-time 2 -o /dev/null http://127.0.0.1:{port}/json/version"
while True:
    remaining = deadline - time.monotonic()
    if remaining <= 0:
        raise RuntimeError(
            f"Chromium did not start in the sandbox; see {PROFILE}/chromium.log there"
        )
    if sandbox.process.exec(probe, timeout=int(remaining) + 1).exit_code == 0:
        return
    time.sleep(0.5)
```

Two bounds keep that wait from hanging. `--max-time 2` stops curl waiting forever on a port that accepts the connection but never answers, and the exec timeout is whatever is left of the 60-second deadline, rounded up to a whole second so the last fraction is still probed and the timeout is never zero. The deadline is checked before every probe, so a Chromium that never comes up ends the run with a pointer to its log instead of spinning.

Once the port answers, the window is up and the desktop is handed to the model.

### 4. How the Driver Maps Toolset Members to the Sandbox

Every member of the computer toolset becomes a Computer Use API call. This is the part to understand if you plan to drive something other than a sketchpad:

| Toolset member | Daytona call |
| --- | --- |
| `screenshot` | `computer_use.screenshot.take_full_screen()` |
| `zoom` | `computer_use.screenshot.take_region(...)`, the crop scaled back up to a full screenshot's size so small detail stays legible |
| `left_click`, `right_click`, `middle_click`, `double_click`, `triple_click` | `computer_use.mouse.click(x, y, button, clicks=..., modifiers=...)` |
| `left_mouse_down` / `left_mouse_up` | `computer_use.mouse.down()` / `computer_use.mouse.up()` |
| `mouse_move`, `cursor_position` | `computer_use.mouse.move(...)` / `mouse.get_position()` |
| `left_click_drag` | `computer_use.mouse.drag(start_x, start_y, end_x, end_y, modifiers=...)` |
| `scroll` (up, down, left, right) | `computer_use.mouse.scroll(x, y, direction, amount, modifiers=...)` |
| `key` | three routes: a named Daytona key goes through `computer_use.keyboard.press(key, modifiers)`; a single character that has no named key, pressed with no modifiers, goes through `computer_use.keyboard.type(char)`; modifier-only chords (`super`, `ctrl+alt`), numpad Enter and the numpad operators, and keysyms with no Daytona key name (`XF86…`) go through the in-sandbox XTest helper |
| `hold_key` | `keyboard.down(...)`, sleep for the requested duration, `keyboard.up(...)` in reverse order |
| `type` | `computer_use.keyboard.type(text)` |
| `wait` | sleeps in your process; nothing is sent to the sandbox |

A few behaviors fall out of that mapping and matter in practice:

- **Modifier-only chords ride along with the action.** Shift-clicking is one `mouse.click` with `modifiers=["shift"]`, not a separate key-down/key-up dance. The drawing task in this example leans on exactly that: shift-click chains into a polyline in `paint.html`. A chord that holds anything other than modifiers takes the XTest route below instead.
- **Horizontal scroll is native.** `scroll_direction` of `left` or `right` goes to `mouse.scroll` like any other direction, which is what lets the model reach swatches that have scrolled off the palette.
- **Exotic keys and non-modifier chords fall back to XTest.** Keysyms the keyboard API doesn't name (`XF86…`, `Print`) and modifier-only chords are sent as raw XTest events by a small helper the driver uploads into the sandbox. So is any click, drag or scroll whose held chord carries a non-modifier key token, and that rule does not depend on whether the key is expressible natively: `left_click` with `text="ctrl+a"` goes to XTest even though `key` with `ctrl+a` is a plain native `keyboard.press("a", ["ctrl"])`. Nothing is installed; the helper talks to libX11/libXtst through `ctypes`.
- **Numpad Enter and the numpad operators take that same XTest route.** `num_enter`, `num_asterisk`, `num_minus`, `num_plus` and `num_slash` are sent as their `KP_*` keysyms, because the native key press emits the wrong characters for them on the daemon current sandboxes run (0.222.1 — `num_enter` arrives as a backtick, `num_plus` as `n`, and so on). The remaining numpad keys — the digits, `num_decimal`, `num_equal` and `num_lock` — were checked byte for byte, are correct natively, and go through the keyboard API.
- **Coordinates are checked, not clamped.** A point outside the screen comes back to the model as an error, so a misread screenshot surfaces instead of silently clicking the edge.
- **Screenshots settle first.** Each screenshot waits out a short delay (0.3 s by default) after the last input, so it shows the effect of what just happened. If the real screen is bigger than `max_screenshot_size`, it is scaled down for the model and the model's coordinates are scaled back up.

:::note[Ahead of the API reference]
The [Computer Use API reference](https://www.daytona.io/docs/en/computer-use.md) documents mouse and keyboard hold, but not everything the driver uses: `clicks` and `modifiers` on `mouse.click`, `modifiers` on `mouse.drag` and `mouse.scroll`, and the `left` and `right` scroll directions, which its Scroll section still describes as rejected. Those options arrive with the same newer computer-use endpoints the platform floor below is about, and they work on any sandbox that clears it. The numpad caveat above is not in the reference either — the keys are supported by the API; it is this daemon build that mis-sends five of them. Where the two pages disagree, the reference is the one that lags.
:::

:::caution[Platform floor for held input]
Native held and extended input — `left_mouse_down`/`left_mouse_up`, `hold_key`, `type`, triple clicks, horizontal scroll, any click, drag or scroll with modifiers, and the numpad keys that go through the keyboard API (the digits, `num_decimal`, `num_equal` and `num_lock`) — needs a sandbox whose Daytona daemon ships the newer computer-use endpoints. The driver settles that once per instance with a deliberately malformed `mouse.down`: a daemon that has the endpoints rejects it with `400` before any input happens, one that does not answers `404`. On `404` every member in that list fails with a clear message telling you to recreate the sandbox on a current version, rather than silently doing nothing. Nothing on the XTest route is subject to the floor, so numpad Enter and the numpad operators keep working either way.

Let the driver create the sandbox, as this example does, and you get Daytona's default snapshot, which clears the floor today: daemon 0.222.1, probe answered `400`. Pinning your own image does not necessarily — a sandbox created from `daytonaio/sandbox:latest` came up on daemon 0.217.0, answered `404`, and refused every member in that list. Prefer the default snapshot unless you have a reason not to.
:::

:::tip[Approval is mandatory for keyboard members]
The Anthropic SDK requires a `confirm` callable whenever `type`, `key` or `hold_key` is enabled. The example passes one that logs the member and its arguments, then returns `True`:

```python
def confirm(context: BetaComputerConfirmContext) -> bool:
    call = json.dumps(context.input.model_dump(exclude_none=True))
    print(f"[confirm] {context.member} {call}")
    return True
```

Blanket approval is fine for a throwaway desktop. Gate it properly the moment the sandbox holds anything you care about.
:::

### 5. Run the Example

```bash
source .env
python draw.py
```

Each line of output is one step of the loop:

```text
sandbox 7f3c…, screen 1280x800
sketchpad open at file:///tmp/claude-draws/paint.html; handing the desktop to the model
[confirm] screenshot {}
[screenshot] {}
[text] I can see the sketchpad. The palette along the bottom…
[confirm] scroll {"coordinate": [640, 720], "scroll_direction": "right", "scroll_amount": 3}
```

The last line totals what the run cost, summed over every assistant turn:

```text
tokens: 412905 input, 6124 output
```

Screenshots dominate the input count: the model takes one after most actions, and each is billed as an image. Cache and server-tool counters are separate fields in the API response and are not folded into these two numbers.

The task prompt is fixed, in `TASK` at the top of `draw.py`: draw a sunset over water, an orange sun, a teal horizon with waves, rose clouds, and a caption. The sandbox is deleted when the run ends.

:::caution[The result is not reproducible]
Model-driven GUI runs are not deterministic. The same prompt will not produce the same picture twice, and the model can misread the screen or miss a swatch entirely. Treat the drawing as a demo of the loop, not as a render you can diff.
:::

### 6. Adapting the Example

**Rewrite the task, keep its two constraints.** Edit `TASK` to draw something else, but keep both rules that make the run work: target colors by their visible label (`orange`, `teal`, `rose`), never by position, because the palette scrolls horizontally and swatch coordinates are not stable; and stay inside the drawing band, roughly 120 to 620 pixels from the top, where the floating toolbar and palette do not intercept clicks.

**Give the model a feedback signal.** `paint.html` shows the selected color's name in the top toolbar, so the model can screenshot and confirm a click landed before it draws. Any app you point this at benefits from the same thing: visible state the model can read back is worth more than a precise click.

**Drive an app that isn't a sketchpad.** Nothing in the loop knows about canvases. Swap `launch_chromium` for whatever starts your GUI on the sandbox's `DISPLAY=:0`, and rewrite the prompt. The toolset mapping in step 4 is unchanged.

**Bring your own sandbox.** Pass a running `Sandbox` to `DaytonaComputer(sandbox)` and the driver drives it without ever stopping or deleting it. With no sandbox, it creates one and deletes it on close — or stops it instead with `on_close="stop"`. Two things follow. A sandbox that outlives the run is no longer throwaway, so replace the blanket `confirm` from step 4 with one that actually inspects `context.member` and its arguments before returning `True`. And a sandbox you built yourself has to clear the platform floor above; a sandbox the driver creates from the default snapshot already does.

**Change the screen size.** `resolution=(width, height)` sets the desktop of a sandbox the driver creates. Remember that the coordinate advice baked into `TASK` is written for 1280x800; if you change the resolution, update the band.

### 7. Key Advantages

- **The desktop is disposable.** Every run gets a fresh sandbox that is deleted at the end, so an agent with a mouse can't reach anything that outlives the task.
- **Your keys never leave your process.** The model loop and the toolset run locally; the sandbox only receives input events and returns screenshots.
- **The full toolset is implemented**, including the members that need held input, so you are not writing per-member workarounds.
- **Errors are legible to the model.** Off-screen coordinates and unsupported platform features come back as plain messages the model can act on, instead of inputs that quietly go nowhere.