Skip to content
← All projects
Computer use
Android prototype

Independent R&D · Source snapshot: 30 September 2026

Perceptron Phone

Operating an existing Android interface requires more than a correct tap: the agent must check whether the intended task was completed.

An agent that operates Android apps from a spoken goal, with checks before consequential actions and before declaring a task complete.

Watch the demo ↓

What this project demonstrates

Changing the planner did not resolve a shared failure on right-filling timer keypads.

  1. Spoken goal
  2. Screen + element tree
  3. Mk1.5 action plan
  4. Android action + checks
  5. Observe again
Architecture from the implementation · Perceptron Phone

Watch the phone demo

54.7 seconds · Edited demonstration · Speed changes are labelled in the recording.

The recording shows the Perceptron interface accepting spoken requests, then interacting with Android apps. Visible examples include setting an alarm, requesting directions and opening Spotify to play Nirvana. This edited demonstration is not an end-to-end benchmark or a measurement of real-time latency.

Open Watch the phone demo video

Scope and contribution

Raoul developed the Android agent loop, accessibility integration, run recording and model-comparison tooling. Perceptron AI supplies Mk1.5, the hosted planning model; Raoul did not train it. Android supplies the accessibility and speech interfaces.

Problem and constraints

A useful phone agent has to do more than identify a button. It must interpret the current state, choose an action, execute it, and check whether the intended change happened. A correct-looking tap can still leave the task unfinished.

The prototype uses Android accessibility rather than a modified operating system. Its view combines screenshots and a numbered element tree. Covered, disabled or vanished controls need different handling from visible, actionable ones.

How the agent works

A spoken instruction starts a run. The app observes the screen and supplies the model with the goal, interaction context and available actions. Mk1.5 chooses an action using the numbered elements; the Android execution layer applies it and captures a new observation.

The loop includes checks before start or confirm actions, a separate completion check, repeat guards, and a stop when the phone is locked. Run recording preserves screenshots and model requests so decisions can be inspected later. These controls are implemented safeguards, not a guarantee that every task is safe or correct.

Evaluation and its limits

Two offline comparisons answer different questions: single-step tasks on saved screens test a requested action; replayed steps test agreement with actions from previously successful Mk1.5 runs. Replay agreement is not fresh end-to-end task success. A different valid path can be scored as a miss.

UI-TARS-1.5-7B and GUI-Owl-1.5-8B-Instruct were evaluated as alternative planners. The open models use screenshot and coordinate formats, while Mk1.5 uses the app’s prompt, tools and element tree. Converted saved-screen trees only approximate the live app’s observations, so the inputs are not equivalent.

The project notes favour retaining Mk1.5 for latency and integration reasons. They also report shared keypad failures. Full benchmark scores are omitted here because the underlying result files and complete run conditions have not been independently checked for this case study.

What failed

On a right-filling timer keypad, the displayed value after entering a digit was mistaken for the completed target time. Replacing the model did not remove that failure. Other recorded limitations include search fields that do not become editable and navigation to the wrong settings category.

Completion remains a separate problem: an agent may announce success too early, or keep acting when the desired result is already visible. A successful tool call is therefore not the same as a successful task.

Decision and next test

Keep Mk1.5 as the current planner and retain the alternative-model evaluation tools. The next useful test is a live comparison from repeatable starting states, including failure recovery, rather than a larger claim based on replay agreement.

This is a prototype, not a published consumer app. The edited demonstration shows example interactions; an independently inspectable evaluation bundle is still needed to assess repeatability and failure rates.

Practical implications

This prototype explores automation where a workflow is exposed through an app interface. It demonstrates observation, action execution and completion checks; it does not establish unattended operation.

A similar engagement would first need a permitted workflow, repeatable starting states, explicit approval boundaries and live recovery tests on the target apps.

Evidence

The implementation and project notes support this account. The supplied demonstration is playable above; the full evaluation result files remain unavailable for this case study.

Discuss a related workflow or integration

Discuss an AI project

Describe the workflow, integration or technical question you want to explore.

Discuss a project

info@genaisolutions.net