There are three ways to control a box. You can use more than one in the same task.
| Mode | Acts on | Use it for | Needs |
|---|---|---|---|
| Screen coordinates | Pixels on the screen | Any visible application, browser chrome, and system prompts | Nothing |
| Accessibility | Native widgets, by role and name | File dialogs, settings panels, installers, and other native forms | A box with accessibility |
| Browser | Page elements, through Chrome DevTools | Web pages | Nothing on host runtimes. On E2B, a template built from the current image. On Vercel, nothing. |
Select a mode
- Is it a web page? Use the browser mode.
- Is it a native window that publishes its widgets? Use the accessibility mode.
- All other cases: use screen coordinates.
Browser chrome (the address bar and tabs), permission prompts, and native file choosers are outside the page. Use screen coordinates for them, or a dedicated tool: upload_file fills a file input with no chooser, and dialog answers alerts and prompts.
Screen coordinates
Take a screenshot, find a point, and send mouse or keyboard input to it.
holm screenshot "$BOX" screen.png
holm mouse "$BOX" click 640 400 left
holm keyboard "$BOX" type "hello"let png = computer.screenshot().await?;
computer.click((640, 400), Button::Left).await?;
computer.type_text("hello").await?;MCP: screenshot, then click, type_text, press_key, drag, and scroll.
Rules:
- Coordinates are device pixels on one screen.
(0, 0)is the top-left corner. - Get the coordinates from the most recent full-size screenshot. A point from an old, cropped, or scaled image goes to the wrong place, and nothing tells you.
- After an action that draws, such as opening a menu, wait until the screen is still (
wait_until_still) before the next screenshot. - A screenshot does not show the pointer. Use
cursorto get its position.
This mode works on all runtimes and with all applications. It is also the easiest to break: a page that moves, a window that opens, or a different screen size changes the correct point.
Accessibility
Native applications publish a tree of widgets through AT-SPI. Each widget has a role, such as push button or text, and a name. This mode finds widgets by name, so it does not depend on pixels.
A box has accessibility by default. You cannot add it later to a box launched with --minimal.
BOX=$(holm new)
holm widget "$BOX" find "Street" --role text
holm widget "$BOX" fill "Street" "12 Bishop Street"
holm widget "$BOX" press "OK"let computer = Computer::builder().launch().await?;
let screen = computer.primary();
let street = NodeQuery {
query: "Street".to_string(),
role: Some("text".to_string()),
..NodeQuery::default()
};
screen.set_node(&street, "12 Bishop Street").await?;MCP: launch_box with accessibility: true, then widget.
Rules:
- A query matches the widget name and also the label next to it, because a form field usually has no name of its own.
pressruns the widget's own action and sends no pointer event. It works on a widget that is covered or out of view.- If the application must receive a real click, use
findto get the widget's rectangle, then click with screen coordinates. - Action names come from the toolkit. GTK calls an action
click, and Qt calls itPress.findlists them. - Some applications and custom widgets publish an incomplete tree. Use screen coordinates for them.
Browser
The browser mode controls Chromium through the Chrome DevTools Protocol (CDP). It acts on page elements, so it continues to work when the window moves or another window covers it.
TAB=$(holm open "$BOX" https://www.selenium.dev/selenium/web/web-form.html)
holm browser "$BOX" snapshot --tab "$TAB"
# @e2 textbox "Text input"
# @e8 combobox "Dropdown (select)"
# @e12 checkbox "Default checkbox"
# @e15 button "Submit"
holm browser "$BOX" fill @e2 "agent"
holm browser "$BOX" select @e8 "Two"
holm browser "$BOX" check @e12
holm browser "$BOX" click @e15
holm browser "$BOX" wait "Received!" --within 10000MCP: open_url, snapshot, then fill_field, dropdown, check, click_element, and wait_for.
Queries
A query can be:
- Visible text
- An accessible name
- An element ID
- A placeholder
- A CSS selector
- A reference from
snapshot, such as@e12
Use a reference after a snapshot, because it names one element. A reference stays the same across snapshots of the same page. A navigation clears all references.
find returns a selector that names exactly one element. Use that selector in the next action, not the text, which can match more than one element.
Add exact when part of the text can match the wrong control.
Use the correct action
| Control | Action |
|---|---|
| Text field, date, color, slider | fill |
| Checkbox, radio | check / uncheck |
| Dropdown | select, deselect, options |
| File input | upload |
| Button, link | click |
fill refuses a dropdown, checkbox, file input, or button, and names the correct action. A native dropdown opens a menu that no screenshot shows and no click can reach, so select is the only way to use one.
Field actions follow an HTML label to its control. They do not guess from nearby text. When a match is ambiguous, the action refuses and names the fields near it.
Wait, do not sleep
After an action that loads a page or fetches data, use wait (wait_for in MCP), not a fixed delay. A fixed delay is either too short or wastes time on each step.
- Wait for an element to appear, or with
gone, to go away. - Give
orthe text that the page shows on failure, such assold out, so the wait stops early. - Use
quiet_msto wait until the page stops changing.
Page and screen together
open opens a new tab and brings it to the front. The screen then shows that tab. Before you mix page actions and screen coordinates, make sure that the correct tab is at the front:
holm browser "$BOX" switch "$TAB"
holm screenshot "$BOX" --tab "$TAB"Other CDP libraries
holm cdp gives an address for Playwright, browser-use, agent-browser, or another CDP library:
export CDP=$(holm cdp "$BOX")
agent-browser --cdp "$(holm cdp "$BOX" --ws)" snapshot -iThe address goes through the server and contains a short-lived token, valid for one hour or for --ttl.
Browser mode on remote runtimes
On Vercel, page tools and holm cdp reach Chromium at the address that Vercel publishes for port 9223, through the same bridge. The server builds the image, so it is current. Daytona and Modal work the same way.
On E2B, page tools and holm cdp reach Chromium at the address that E2B publishes for port 9223. A bridge in the box answers only requests that carry a secret made for that box, so the address alone does not open the browser. Build the template from the current image: an older template does not have the bridge, and Chromium refuses the requests.
Each CDP call goes to the vendor and back, so a page tool takes longer than on a host runtime.
When a server takes a box back after a restart, it restarts the bridge in the box with a new secret, so the secret of an earlier server no longer opens the browser. Page tools then work again on the box. Screenshots, input, clipboard, files, commands, and human control continue to work.
Compare the modes
| Screen coordinates | Accessibility | Browser | |
|---|---|---|---|
| Works on | Anything visible | Native applications that publish widgets | Web pages |
| Survives layout changes | No | Yes | Yes |
| Reaches covered or scrolled elements | No | Yes | Yes |
| Real pointer events | Yes | No (press) |
Yes |
| Setup | None | accessibility at launch |
None; on E2B, a template built from the current image |