Skip to main content
The mobile driver controls the phone inside a session. Its reliable loop is:
  1. Observe the current screen.
  2. Point at a target with a locator, by visible text or a description.
  3. Tap, type, swipe, or press a key.
  4. Observe or wait for the expected result.
A locator action finds its target on the screen at that moment and waits until it can act, up to 5 seconds by default, so most steps need no separate wait. Start with a dedicated-phone session, then use the interface that fits your program.
Phone control uses the session’s Device Control Protocol WebSocket. The CLI and SDK drivers expose these actions; they are not REST endpoints.

The basics

Find

Locate something by its text or by describing it.

Tap

Tap a target or press and hold.

Type

Type into a field.

Swipe

Swipe, scroll, or drag.

Keys

Press Enter to submit a form or search.

Observe

Read the whole screen at once.

Wait

Wait for the screen to change before acting.

Screenshot

Capture the current pixels as a PNG.

Semantic targets and coordinates

Use text or a natural-language query for a tappable target whenever possible. Coordinates are appropriate for gestures that move the whole screen or for a fixed point you derived from the current observation. Only tap, long-press, and swipe have raw coordinate forms. Finding, waiting, typing, key presses, observation, and screenshots use their own typed inputs. For complete Python signatures and the corresponding Go methods, see the Driver and Locator references.