cklxx/laya-browser
laya-browser — laya fine-tuned as a browser-agent decision head (drop-in replacement for TypeSafe Jev)
laya (convaiinnovations/laya) is a non-autoregressive "System 1" decision model: one bidirectional encoder pass answers several typed questions (choice / score / noul) with calibrated probabilities, no text generation. This repo fine-tunes it into the decision head of browser-use/jev-ultrafast, whose /v1/systemone request is exactly laya's predict(state, questions): every step, one forward pass (~22 ms at full GPU clock) picks the operation (CLICK / TYPETEXT / SELECT / PRESSENTER / SCROLL / DONE …) and its target element. Trained and evaluated locally on one RTX 4070 Ti SUPER (16 GB), no paid API; a local Qwen3-8B-AWQ only writes the text that TYPE_TEXT types and picks dropdown values.
One model, at the repo root (v19s). Earlier checkpoints are in the commit history. This repo is updated only when a new version is clearly better.
Results (v19s, mmBERT-base 322M)
All numbers below were measured with the harness in code/jev-ultrafast.patch; the v17s column was measured with the harness it shipped with, so part of the gain is the harness (see "What changed").
- Suite C is the honest headline: real multi-step tasks on domains absent from every training source. v19s reaches ~20–26 %; single runs vary by ±4 tasks (site load times, popups), so treat differences smaller than that as noise.
- Suite A lost
books-open-book(0/3), gained nothing else; suite B unchanged.
Latency. One decision is 22–27 ms on an RTX 4070 Ti SUPER while the GPU is at full clock (requests back to back). An agent waits seconds between decisions for pages to load, and the GPU drops to idle clocks (P5/P8, 210–850 MHz) within a second or two: the next decision then takes 75–270 ms (the demo above shows these live numbers). Locking the minimum SM clock removes that — sudo nvidia-smi -lgc 2100,3135 (undo: sudo nvidia-smi -rgc; costs some idle power). The first request of each input-length bucket also compiles a kernel once (~100–500 ms); warm up with a few requests after starting the server.
Use
transformers (the model ships its own code; only torch + transformers are needed):
from transformers import AutoModel
m = AutoModel.from_pretrained("cklxx/laya-browser", trust_remote_code=True).to("cuda") # CPU works too
page = {"url": ..., "title": ..., "text": ..., # the page: interactive elements + visible text
"actions": [{"id": "e1", "kind": "fill", "node": 1, "label": "Search packages", "role": "searchbox"},
{"id": "e2", "kind": "click", "node": 2, "label": "argo", "role": "link"}, ...]}
d = m.decide(page, goal="Search packages for 'json' and open the package 'argo'.", history=[])
d["operation"], d["action"], d["confidence"] # "CLICK", {"id": "e2", ...}, 0.88decide builds the request exactly in the format the model was trained on (instructions, compact option strings, form-field summary, 1200 chars of page text, split of choices wider than 60) and runs one forward pass. m.systemone(body) answers a jev-ultrafast /v1/systemone request body. The same checkpoint also loads with the upstream laya package (laya.load("cklxx/laya-browser")); laya_browser.py in this repo wraps it the same way and adds a server:
As a TypeSafe replacement for [browser-use/jev-ultrafast](https://github.com/browser-use/jev-ultrafast):
python laya_browser.py serve --port 8791 # --model <local dir> to use a downloaded copy
TYPESAFE_BASE_URL=http://127.0.0.1:8791 TYPESAFE_API_KEY=local <run jev-ultrafast as usual>It answers jev's /v1/systemone requests identically to the evaluation server (checked on 21 real steps: same operation and target on all 21). The harness improvements behind the suite numbers are in code/jev-ultrafast.patch (apply to jev-ultrafast 1231850); the full evaluation / training setup is in code/ (uv sync --extra fast, code/verify.py).
What changed since v17s
Harness (code/jev-ultrafast.patch, all measured case by case on failures):
- wait for the navigation an Enter / click starts before observing (the agent used to see the old page and think Enter did nothing);
- an "undo guard": for 3 steps after a filter / sort / radio click that changed the page, that control (and its "remove filter" chip, and other options of the same dropdown) is not offered again — the agent used to toggle filters on and off until its budget ran out;
- the text model gets each field's placeholder / type / pattern (dates in
MM/DD/YYYYfields) and picks the value of a dropdown the policy chose (ages, times, countries that differ by a digit); - options scrolled out of view inside an open list are offered and scrolled into view before clicking.
Data (v19s = mmBERT-base v17s continued, full fine-tune):
- WebChain (CC-BY-4.0): 3,000 human trajectories on real sites, 10.9k steps after dropping unlabeled / duplicate-label targets; the gold element is recovered exactly from each step's DOM snapshot via its CSS selector (
code/finetune/convert_webchain.py); - webgym grown to 7 task kinds, plus DAgger on webgym (the model drives, the scripted expert labels every visited state);
- format v5 (above).
What still fails
- Stopping too early is ~60 % of real-site failures: after a search the model often says DONE on the results page when the task asks to open a result or a sub-page ("find the recipe page of X", "open its episode list"). More DONE data (Go-Browse), cost-sensitive training, a noul completion head, goal-contrast twins and a run-time sub-goal planner were all tried; none moved suite C beyond noise (details in
results/v19s/). A 322M single-pass policy does not reliably read these goal distinctions. - Pages whose target is several screens down behind many links; sites that block headless Chromium (roughly a third of the Online-Mind2Web sites); logins (jev hides password fields by design).
Changelog
Files
model.safetensors, encoder/, tokenizer/, rl_agent_config.json the model (v19s; a laya checkpoint dir)
modeling_laya_browser.py, config.json transformers remote code (AutoModel + trust_remote_code)
laya_browser.py the same on top of the laya package, plus a TypeSafe-compatible server
code/ server, suites A/B/C, Online-Mind2Web runner + judge, finetune pipeline (webgym, WebChain / Go-Browse converters),
TileLang kernels, jev-ultrafast patch, verify.py
results/ suite JSONs, traces and logs behind the numbers above (results/v19s/, older versions in their folders)
assets/ demo videoLicense
Apache-2.0, same as laya. Mind2Web, NNetNav and WebChain are used under their own licenses for training only.
