Skip to main content
Argus is Axilio’s vision model, and it is how the driver sees the screen. It runs two families of models. The Axilio models read text, icons, and layout for get_by_text() and observe(). The vision-language models understand a description for locator(query=...). The public catalog lists both families with live pricing.
The CLI does not have a model-catalog command. Use an SDK or the public REST operation, then pass the chosen engine or model to axilio phone.

Axilio models

The Axilio line comes in two tiers: Lite is free, and Pro reads harder screens such as small fonts and dense layouts for a per-call price. You do not pass these model IDs to the driver. Pick the tier with an OCR-engine option:
phone find-text and phone wait-for use their default OCR engine and do not accept an engine override.
To use a tier for a whole SDK driver instead of every call, set it once:
The IDs below are what you see on usage rows and in the pricing list:

Vision-language models

Describing something in plain English, such as locator(query="the heart icon"), runs a vision-language model. So does refining a text locator with nth, within, has, or filter. Pass a catalog ID to choose one:
Python and Go can also set a driver default:
Leave model= off and Argus picks a sensible default. An id that isn’t in the list below is rejected before anything is billed. Vision-language calls are billed per token:
Prices are pulled live, so they’re never stale. See Usage to track what you actually spend.

Calling Argus directly

The driver runs Argus for you on the live screen. If you have your own image to analyze, call it directly with client.argus — the same models, on any image. See the Client reference for the SDK methods (locate, detect, list_models), or the Vision API reference for the REST endpoints.

Next steps

Find

Where the engine and model options plug in.

Usage

See what your vision calls cost.