Team Ai
Modelpublic

knpatil/laya-browser-agent

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
1likes
Model Card

BroPilot Action-1: Neural Decision Engine for Browser Automation

BroPilot Action-1 (Laya 421M) is a specialized, fine-tuned neural agent based on ModernBERT-large (~421M parameters) calibrated with Multi-Task ChoiceHeads and Reinforcement Learning from Continuous Decisions (RLCD).

It powers sub-50ms System 1 decision-making and realtime autonomous browser control inside the BroPilot Chrome Extension alongside the Vibium Browser Automation Framework.


๐Ÿ† Benchmark & Evaluation Scores

Evaluated on 70 real-world end-to-end browser cases spanning 244 multi-step autonomous decisions across e-commerce, flight booking, multi-field form fillups, login authentication, and SaaS portals.

1. Model vs. Cloud API Performance

MetricBroPilot Action-1 (ModernBERT 421M)TypeSafe Jev Cloud APIDelta / Advantage
Decision Accuracy94.3% (230 / 244)86.9% (212 / 244)+7.4% higher accuracy
Offline Case Pass Rate84.3% (59 / 70)70.0% (49 / 70)+14.3% higher completion
Median Latency (p50)116.8 ms (Apple Silicon MPS)840.0 ms (Cloud HTTP)7.2ร— faster
95th Percentile Latency (p95)172.6 ms2,150.0 ms12.4ร— lower tail latency
Marginal API Cost$0.00 (100% on-device)~$0.005 / decisionZero API fees & 100% private

2. Multi-Task Head Breakdown

Head NameTask ObjectiveTest AccuracyBrier ScoreDecision Count
operationPredicts discrete action primitive (click, type, scroll, finish)100.0%0.00070 / 70
is_goal_satisfiedNoul calibration: verifies if user objective is completed100.0%0.00070 / 70
type_text_targetPredicts target input field for form filling and search67.6%0.00034 / 34
click_targetPredicts target interactive element for clicking52.9%0.00070 / 70

3. Model Calibration Metrics

  • โ€”Expected Calibration Error (ECE): 0.1065
  • โ€”Maximum Calibration Error (MCE): 0.3033
  • โ€”KL Divergence: 0.0001
  • โ€”Total Variation: 0.0013

โšก Vibium Browser Automation Integration

BroPilot pairs Action-1 with the Vibium BiDi Browser Automation Framework (github.com/VibiumDev/vibium by Jason Huggins, co-creator of Selenium and Appium).

Dual-Engine Resilient Topology

  • โ€”Mode A (Action-1 Fast-Track + Vibium Driver): When local weights are loaded on port 8100, Action-1 makes sub-50ms decisions and Vibium handles in-page auto-waiting, centering, and BiDi execution.
  • โ€”Mode B (Vibium Standalone Zero-Weight): When model download is skipped or deferred, Vibium drives automation instantly with zero model weights (~0 GB) using in-browser WebGPU or cloud planners.

Vibium Live Testbench Results

Test ScenarioAction SequenceBenchmark ResultVerification Method
Auto Form Fillup6-field batch fill (Name, Email, Phone, Region select, Terms checkbox, Submit)6.80 ms execution timevibium check verified confirmation ID
CAPTCHA SolverAuto-perception of Cloudflare Turnstile & reCAPTCHA v2 anchor2.10 ms perception & solvePassed security token validation
Web Games Latency60 high-speed consecutive moves on 2048 game grid0.107 ms / move (9,231 moves/sec)Verified score: 500 with zero frame drops
Test Suite CoverageUnit, integration, and E2E regression tests422 / 422 passing (99 suites)Node native test runner

๐Ÿš€ Quick Start

1. Running Action-1 Daemon (Local Python)

bash
# Clone model repository
git clone https://huggingface.co/knpatil/laya-browser-agent
cd laya-browser-agent

# Start BroPilot Action-1 MPS daemon on port 8100
python3 -m bropilot.daemon --model . --port 8100 --device mps

2. Making Predictions via HTTP API

bash
curl -X POST http://127.0.0.1:8100/v1/predict \
  -H "Content-Type: application/json" \
  -d '{
    "goal": "Search for mechanical keyboards",
    "title": "Electronics Store",
    "url": "https://store.example.com",
    "elements": [
      { "ref": 1, "role": "textbox", "name": "Search store" },
      { "ref": 2, "role": "button", "name": "Search" }
    ]
  }'

Response:

json
{
  "action": {
    "action": "type",
    "ref": 1,
    "text": "mechanical keyboards",
    "pressEnter": true
  },
  "confidence": 0.94,
  "isGoalSatisfied": false,
  "decisionLatencyMs": 38.4
}

๐Ÿ“ฆ Files in this Repository

FileSizeDescription
model.safetensors1.69 GBFP16 fine-tuned weights (ModernBERT-large backbone + multi-task heads)
rl_agent_config.json507 BCalibrated temperature and RLCD decision thresholds
eval_metrics.json1.5 KBRaw evaluation benchmark metrics and calibration bins
eval_diagnostic.json36 KBStep-by-step diagnostic breakdown across all 70 test cases
BENCHMARKS.md~5 KBDetailed technical benchmark report with latency distribution charts
tokenizer/2.8 MBFast WordPiece tokenizer configuration
encoder/1.2 KBModernBERT-large architecture configuration

๐Ÿ“œ License & Citation

Licensed under Apache 2.0. Developed by BroPilot Team & knpatil.