Let Claude Draw on a Daytona Sandbox Desktop
This guide walks through the claude-draws example: a Daytona sandbox boots a desktop, a self-contained HTML sketchpad opens in a visible Chromium window, and Claude is handed the keyboard and mouse through the Anthropic SDK’s computer toolset (computer_toolset_20260801).
Nothing about the app is special-cased for the model. There is no DOM access, no accessibility tree, no scripted click targets. Claude takes a screenshot, finds the palette, scrolls it sideways, clicks a swatch by its visible label, and drags to draw — the same moves a person would make. The bridge between the model’s tool calls and the sandbox is DaytonaComputer from the daytona-claude-toolsets package, which turns each member of the toolset into a call against Daytona’s Computer Use API.
1. Workflow Overview
Section titled “1. Workflow Overview”Your process Daytona sandbox────────────── ───────────────DaytonaComputer(...) ──────────▶ create sandbox, start desktop (1280x800)upload paint.html ──────────▶ /tmp/claude-draws/paint.htmllaunch Chromium (--app=...) ──────────▶ visible window, no tab strip or omnibox │client.beta.messages.tool_runner │ model ⇄ computer toolset │ screenshot ───────────────────┼───▶ screenshot API scroll / click / drag ────────┼───▶ mouse API key / type ───────────────────┼───▶ keyboard API ▼close() ──────────▶ sandbox deletedThe model loop, your API keys, and the toolset all stay in your process. The sandbox only ever sees input events and hands back PNGs.
2. Project Setup
Section titled “2. Project Setup”Requirements
Section titled “Requirements”- Python 3.10 or higher on your machine.
- A Daytona account and API key.
- An Anthropic API key.
Clone the Repository
Section titled “Clone the Repository”Clone the Daytona guides repository and set up the example:
git clone https://github.com/daytona/guides.gitcd guides/python/computer-use/claude-drawspython3 -m venv venvsource venv/bin/activate # On Windows: venv\Scripts\activatepip install -e .pip install -e . pulls in daytona-claude-toolsets (which supplies DaytonaComputer) and an anthropic release new enough to ship anthropic.tools.computer. daytona-claude-toolsets in turn requires the daytona SDK 0.223.x — the first release whose Python client exposes the mouse and keyboard hold calls the driver routes through.
Configure API Keys
Section titled “Configure API Keys”cp .env.example .envFill in:
DAYTONA_API_KEY: required. Get it from the Daytona Dashboard.ANTHROPIC_API_KEY: required, for the model loop.MODEL: optional, the model the drawing loop uses. Defaults toclaude-sonnet-5-5.
3. Opening the Sketchpad
Section titled “3. Opening the Sketchpad”DaytonaComputer is a context manager. Entering it creates the sandbox, starts its desktop, waits for a display to report a size, and exposes the live Sandbox object so you can set the screen up before the model ever looks at it:
with DaytonaComputer(confirm=confirm, resolution=(WIDTH, HEIGHT)) as computer: print(f"sandbox {computer.sandbox.id}, screen {computer.width}x{computer.height}") upload_sketchpad(computer.sandbox) launch_chromium(computer.sandbox)paint.html is uploaded with the sandbox filesystem API, then opened in Chromium in app mode (--app=file://...), so the window is nothing but the sketchpad — no tab strip, no omnibox, no first-run bubbles to confuse the model.
One detail is worth copying into your own runs. Chromium is a foreground GUI process that never returns, so launching it with process.exec would block until the browser exits. The example starts it through an async process session instead, then polls Chromium’s loopback debugging port purely as a readiness probe:
sandbox.process.create_session(SESSION_ID)sandbox.process.execute_session_command( SESSION_ID, SessionExecuteRequest(command=command, run_async=True))deadline = time.monotonic() + 60probe = f"curl -sf --max-time 2 -o /dev/null http://127.0.0.1:{port}/json/version"while True: remaining = deadline - time.monotonic() if remaining <= 0: raise RuntimeError( f"Chromium did not start in the sandbox; see {PROFILE}/chromium.log there" ) if sandbox.process.exec(probe, timeout=int(remaining) + 1).exit_code == 0: return time.sleep(0.5)Two bounds keep that wait from hanging. --max-time 2 stops curl waiting forever on a port that accepts the connection but never answers, and the exec timeout is whatever is left of the 60-second deadline, rounded up to a whole second so the last fraction is still probed and the timeout is never zero. The deadline is checked before every probe, so a Chromium that never comes up ends the run with a pointer to its log instead of spinning.
Once the port answers, the window is up and the desktop is handed to the model.
4. How the Driver Maps Toolset Members to the Sandbox
Section titled “4. How the Driver Maps Toolset Members to the Sandbox”Every member of the computer toolset becomes a Computer Use API call. This is the part to understand if you plan to drive something other than a sketchpad:
| Toolset member | Daytona call |
|---|---|
screenshot | computer_use.screenshot.take_full_screen() |
zoom | computer_use.screenshot.take_region(...), the crop scaled back up to a full screenshot’s size so small detail stays legible |
left_click, right_click, middle_click, double_click, triple_click | computer_use.mouse.click(x, y, button, clicks=..., modifiers=...) |
left_mouse_down / left_mouse_up | computer_use.mouse.down() / computer_use.mouse.up() |
mouse_move, cursor_position | computer_use.mouse.move(...) / mouse.get_position() |
left_click_drag | computer_use.mouse.drag(start_x, start_y, end_x, end_y, modifiers=...) |
scroll (up, down, left, right) | computer_use.mouse.scroll(x, y, direction, amount, modifiers=...) |
key | three routes: a named Daytona key goes through computer_use.keyboard.press(key, modifiers); a single character that has no named key, pressed with no modifiers, goes through computer_use.keyboard.type(char); modifier-only chords (super, ctrl+alt), numpad Enter and the numpad operators, and keysyms with no Daytona key name (XF86…) go through the in-sandbox XTest helper |
hold_key | keyboard.down(...), sleep for the requested duration, keyboard.up(...) in reverse order |
type | computer_use.keyboard.type(text) |
wait | sleeps in your process; nothing is sent to the sandbox |
A few behaviors fall out of that mapping and matter in practice:
- Modifier-only chords ride along with the action. Shift-clicking is one
mouse.clickwithmodifiers=["shift"], not a separate key-down/key-up dance. The drawing task in this example leans on exactly that: shift-click chains into a polyline inpaint.html. A chord that holds anything other than modifiers takes the XTest route below instead. - Horizontal scroll is native.
scroll_directionofleftorrightgoes tomouse.scrolllike any other direction, which is what lets the model reach swatches that have scrolled off the palette. - Exotic keys and non-modifier chords fall back to XTest. Keysyms the keyboard API doesn’t name (
XF86…,Print) and modifier-only chords are sent as raw XTest events by a small helper the driver uploads into the sandbox. So is any click, drag or scroll whose held chord carries a non-modifier key token, and that rule does not depend on whether the key is expressible natively:left_clickwithtext="ctrl+a"goes to XTest even thoughkeywithctrl+ais a plain nativekeyboard.press("a", ["ctrl"]). Nothing is installed; the helper talks to libX11/libXtst throughctypes. - Numpad Enter and the numpad operators take that same XTest route.
num_enter,num_asterisk,num_minus,num_plusandnum_slashare sent as theirKP_*keysyms, because the native key press emits the wrong characters for them on the daemon current sandboxes run (0.222.1 —num_enterarrives as a backtick,num_plusasn, and so on). The remaining numpad keys — the digits,num_decimal,num_equalandnum_lock— were checked byte for byte, are correct natively, and go through the keyboard API. - Coordinates are checked, not clamped. A point outside the screen comes back to the model as an error, so a misread screenshot surfaces instead of silently clicking the edge.
- Screenshots settle first. Each screenshot waits out a short delay (0.3 s by default) after the last input, so it shows the effect of what just happened. If the real screen is bigger than
max_screenshot_size, it is scaled down for the model and the model’s coordinates are scaled back up.
5. Run the Example
Section titled “5. Run the Example”source .envpython draw.pyEach line of output is one step of the loop:
sandbox 7f3c…, screen 1280x800sketchpad open at file:///tmp/claude-draws/paint.html; handing the desktop to the model[confirm] screenshot {}[screenshot] {}[text] I can see the sketchpad. The palette along the bottom…[confirm] scroll {"coordinate": [640, 720], "scroll_direction": "right", "scroll_amount": 3}The last line totals what the run cost, summed over every assistant turn:
tokens: 412905 input, 6124 outputScreenshots dominate the input count: the model takes one after most actions, and each is billed as an image. Cache and server-tool counters are separate fields in the API response and are not folded into these two numbers.
The task prompt is fixed, in TASK at the top of draw.py: draw a sunset over water, an orange sun, a teal horizon with waves, rose clouds, and a caption. The sandbox is deleted when the run ends.
6. Adapting the Example
Section titled “6. Adapting the Example”Rewrite the task, keep its two constraints. Edit TASK to draw something else, but keep both rules that make the run work: target colors by their visible label (orange, teal, rose), never by position, because the palette scrolls horizontally and swatch coordinates are not stable; and stay inside the drawing band, roughly 120 to 620 pixels from the top, where the floating toolbar and palette do not intercept clicks.
Give the model a feedback signal. paint.html shows the selected color’s name in the top toolbar, so the model can screenshot and confirm a click landed before it draws. Any app you point this at benefits from the same thing: visible state the model can read back is worth more than a precise click.
Drive an app that isn’t a sketchpad. Nothing in the loop knows about canvases. Swap launch_chromium for whatever starts your GUI on the sandbox’s DISPLAY=:0, and rewrite the prompt. The toolset mapping in step 4 is unchanged.
Bring your own sandbox. Pass a running Sandbox to DaytonaComputer(sandbox) and the driver drives it without ever stopping or deleting it. With no sandbox, it creates one and deletes it on close — or stops it instead with on_close="stop". Two things follow. A sandbox that outlives the run is no longer throwaway, so replace the blanket confirm from step 4 with one that actually inspects context.member and its arguments before returning True. And a sandbox you built yourself has to clear the platform floor above; a sandbox the driver creates from the default snapshot already does.
Change the screen size. resolution=(width, height) sets the desktop of a sandbox the driver creates. Remember that the coordinate advice baked into TASK is written for 1280x800; if you change the resolution, update the band.
7. Key Advantages
Section titled “7. Key Advantages”- The desktop is disposable. Every run gets a fresh sandbox that is deleted at the end, so an agent with a mouse can’t reach anything that outlives the task.
- Your keys never leave your process. The model loop and the toolset run locally; the sandbox only receives input events and returns screenshots.
- The full toolset is implemented, including the members that need held input, so you are not writing per-member workarounds.
- Errors are legible to the model. Off-screen coordinates and unsupported platform features come back as plain messages the model can act on, instead of inputs that quietly go nowhere.