Team Ai
Modelpublic

thealper2/smollm2-135m-android-control

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes208downloads
Model Card

smollm2-135m-android-control

LoRA SFT of HuggingFaceTB/SmolLM2-135M-Instruct as a text-only Android UI action policy. Input: instruction + textual UI element list + action history. Output: one action as a single-line JSON object. The model does not consume images.

Output schema

action_typefields
clicktarget (element index) or x,y (0-1000)
long_presstarget or x,y (0-1000)
scrolldirection ∈ up/down/left/right
input_texttext
open_appapp_name
wait / navigateback / navigatehome-

Normalized coordinates: x = round(px / screen_width * 1000), same for y. Element bounds in the prompt use the same 0-1000 grid.

Prompt format

Chat template of the base tokenizer, system + user turn, greedy decoding.

System:

text
You are an Android UI control agent. Given the instruction and the current screen, output the single next action as raw JSON and nothing else.
Actions: {"action_type":"click","target":N} | {"action_type":"long_press","target":N} | {"action_type":"input_text","text":"..."} | {"action_type":"scroll","direction":"up|down|left|right"} | {"action_type":"open_app","app_name":"..."} | {"action_type":"wait"} | {"action_type":"navigate_back"} | {"action_type":"navigate_home"}
N is an element index from the UI list. If no element fits, use {"action_type":"click","x":X,"y":Y} with X and Y in 0-1000.

User:

text
Goal: <episode goal>
Instruction: <step instruction>
Current Android UI:
screen 1080x2400 app=com.android.settings
[0] EditText "Search settings" [56,108,944,167] click,edit
[1] TextView "Network & internet" [56,200,944,267] click
...
Previous actions:
1. open_app "Settings"
Return exactly one action.

Element line: [index] class "label" [left,top,right,bottom] flags. Label = text > content-description > hint. History entries reference element labels, not indices.

Data

  • —Source: OfficerChul/Android-Control-84k @ 0248027f747c9d57bd09c14e8f044f9a8103dddd (step instructions, actions, screenshot names).
  • —UI text: accessibility trees from the original AndroidControl TFRecords (gs://gresearch/android_control), joined by episode id and step.
  • —Element selection: visible, enabled, labelled or interactive nodes; status bar removed; top 32 by score, reading order; labels truncated to 40 chars.
  • —Click/long-press target: smallest kept element containing the gold point; otherwise normalized coordinates.
  • —History: last 5 actions.
  • —Upstream rows (train+test): 83,848; converted: 83,848; discarded: 0
  • —Rows with an accessibility tree: 83,848; without: 0
  • —Position encoding: coords 5,919, none 31,763, target 46,166
  • —Token budget: 1280 (prompt+completion); rows with elements dropped to fit: 15,889
  • —Train rows from episodes that also occur in the upstream test split removed: 5,642
splitrowsepisodes
test904851
train75,19514,023
validation2,107400

Training

parametervalue
methodLoRA
r / alpha / dropout32 / 64 / 0.05
target modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
trainable params9,768,960 (6.7706%)
total params144,283,968
losscompletion-only (assistant action tokens)
epochs2
batch size × grad accum8 × 4
learning rate / scheduler0.0002 / cosine
warmup ratio / weight decay0.03 / 0.01
max length1280
precisionbf16
seed42
hardwareNVIDIA GeForce RTX 5060 Ti
softwaretorch 2.11.0+cu128, transformers 5.17.0, trl 0.24.0, peft 0.18.1
metricvalue
train_loss0.1554
eval_loss0.113
train_runtime4955.4101

Evaluation

Split: test (n=904), identical prompts for both models, greedy decoding. Strict JSON = the completion is exactly one schema-valid JSON object. Click target accuracy requires the correct action type and index. Coordinate accuracy: L∞ distance ≤ 50 on the 0-1000 grid.

Metricbasefinetuned
JSON validity (strict)0.0000.999
JSON validity (lenient)0.0000.999
Exact action accuracy0.0000.695
Action type accuracy0.0000.883
Click target accuracy0.0000.488
Coordinate accuracy (<=5%)0.0000.122
Input text accuracy0.0000.792
Scroll accuracy0.0000.912
Open app accuracy0.0000.824

Per-action exact match:

Actionnbasefinetuned
click1250.0000.592
input_text1250.0000.792
long_press1250.0000.264
navigate_back1250.0000.976
navigate_home290.0000.000
open_app1250.0000.824
scroll1250.0000.912
wait1250.0000.664

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "thealper2/smollm2-135m-android-control"
model = AutoModelForCausalLM.from_pretrained(repo)
tok = AutoTokenizer.from_pretrained(repo)

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},  # see "Prompt format"
    {"role": "user", "content": user_prompt},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt", add_special_tokens=False)
out = model.generate(**inputs, max_new_tokens=48, do_sample=False)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# {"action_type":"click","target":1}

Limitations

  • —Text-only policy: no pixels. Elements absent from the accessibility tree (canvas, games, some WebViews) are not visible to it.
  • —Target indices are only meaningful for the element list they were predicted on; the same element selection rules must be used at inference.
  • —Training data has no navigate_home examples and only 42 long_press examples; the test split is class-balanced, so those classes dominate the error.
  • —open_app emits an app display name; mapping to a package is done by the executor.
  • —Single-step accuracy on offline data does not measure closed-loop task success.