Skip to main content
The driver is what client.session() hands you inside a session. It’s the object every action runs on: finding, tapping, typing, swiping, and reading the screen. Methods that point at something return a Locator; reading the whole screen returns a Screen.
The detailed signatures below use Python. The Go SDK exposes the same DCP surface with Go naming and explicit duration arguments; see the method map at the end. Live phone actions are not REST operations.

Connection and capabilities

driver.handshake()

Return the executor’s protocol version, device descriptor, domains, and exact capability set. Gate optional behavior on these advertised capabilities.
Returns HandshakeResult.

driver.device_info()

Return the phone’s static DCP descriptor.
Returns DeviceInfo.

driver.close()

Close the underlying control transport. client.session() calls this and deallocates its phone automatically when the context exits.

Locating

driver.get_by_text()

Point at something by its visible text, read by OCR. Nothing is sent until you call an action on the result.
Returns Locator.

driver.locator()

Point at something by describing it in plain English, for when there’s no text to match (an icon, an image). A description runs a vision-language model, so it costs more than text.
Returns Locator.

Tapping

driver.tap()

Tap an exact point on the screen. To tap something you located, call locator.tap() instead.
Returns None. coords is a {"x": int, "y": int} dict.

driver.long_press()

Press and hold an exact point, for example to open a context menu. To hold a located target, read its position with locator.bounding_box().
Returns None.

Typing

driver.type_text()

Type into the field that’s already selected. To tap a field and type in one call, use locator.fill(). To press Enter (submit / search), use driver.key_press().
Returns None.

Swiping

driver.swipe()

Swipe between two points. This is also how you scroll: swipe up (end higher than start) to move down a list. To drag one located target to another, see Swipe.
Returns None.

Buttons

driver.key_press()

Press a named key. Key.ENTER submits a form or fires the on-screen keyboard’s Go / Search.
Returns None. key is a string; see Key for the constants.

driver.press()

Press a named key against whatever has focus, through the same path as a locator action. To focus a specific target first, use locator.press().
Returns LocatorResult with no bounds, since nothing was located.

Reading the screen

driver.observe()

Read everything on the screen at once, all text and icons, as a queryable snapshot you can search without going back to the phone.
Returns Screen.

driver.screenshot()

Take a picture of the screen as PNG image data, to save, log, or debug a run.
Returns bytes: raw PNG data.

Go method map

Go constructs a driver with mobile.ConnectRemote(controlURL, options...). Use mobile.WithDefaultOCREngine, mobile.WithDefaultModel, and mobile.WithOpenTimeout for connection-wide defaults. Observe, Handshake, and DeviceInfo take mobile.WithOCREngine. Locators take the options listed in the Locator reference. Go requires gesture duration arguments where Python supplies defaults.

See also

Locator

What get_by_text() and locator() return, and the actions on it.

Screen

What observe() returns.

Client

Opening a session and calling Argus directly.

Key

Every device-button constant.