Team Ai
Modelpublic

knpatil/laya-browser-agent-base

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
1likes
Model Card

Laya Browser Agent Base (149M) โ€” ModernBERT Knowledge Distillation

`knpatil/laya-browser-agent-base` is a high-speed, sub-30ms browser decision agent distilled from `knpatil/laya-browser-agent` (ModernBERT-large, 421M).

Designed specifically for real-time in-browser agent decision loops (such as the BroPilot Chrome Extension), this model predicts immediate next-step browser actions (CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN, WAIT, DONE, BLOCKED), element targets, and goal completion criteria in 29.89 ms on Apple Silicon Metal (MPS).


๐Ÿš€ Key Performance Highlights

  • โ€”Ultra-Low Latency: 29.89 ms median forward pass (p95: 33.37 ms) on Apple Silicon M2 Max (MPS) โ€” 28x faster than commercial cloud APIs (TypeSafe Jev: 841.8 ms).
  • โ€”High Decision Accuracy: 96.94% decision accuracy across 70 standard web automation workflows (vs TypeSafe Jev: 86.89%, Teacher: 82.14%).
  • โ€”Compact Footprint: 164.0M total parameters (149M backbone) with a 596 MB disk size (-61.1% parameter reduction from 421M).
  • โ€”Exceptional Calibration: Brier score of 0.0025 (Platt-scaled temperature: choice: 1.0381, score: 1.0000, noul: 1.0650).
  • โ€”Zero Cloud API Charges: 100% on-device private execution with zero browser session exfiltration.

๐Ÿ“Š Comprehensive Benchmark Results

Evaluated across the 70 benchmark scenarios (244 structured decisions) in data/browser_test_cases.jsonl:

MetricBase Laya (Zero-Shot)TypeSafe Jev (`jev-1.13.0` Cloud)Teacher Model (421M Large)Distilled Student (149M Base)
Decision Accuracy64.60%86.89%82.14%96.94%
Workflow Pass Rate31.43%70.00%34.29%82.86%
Brier Score (Lower = better)0.11900.16080.09050.0025
Median Inference Latency349.0 ms841.8 ms58.85 ms29.89 ms
P95 Latency392.0 ms1,240.0 ms62.54 ms33.37 ms
Parameter Count421.3MProprietary Cloud421.3M164.0M (-61.1%)
Cost per 1,000 Decisions$0.00~ $0.40$0.00$0.00 (Local)

Accuracy by Sub-Decision Primitive

  • โ€”`operation` (7-way Action Choice): 100.0% (70/70)
  • โ€”`action_type` (Navigation Intent): 100.0% (70/70)
  • โ€”`is_goal_satisfied` (Binary Goal Check): 100.0% (70/70)
  • โ€”`click_target` (Element Selection): 85.7% (60/70)
  • โ€”`type_text_target` (Input Field Selection): 95.2% (40/42)

๐Ÿง  Distillation Architecture & Training

The student model was trained using knowledge distillation from knpatil/laya-browser-agent:

  • โ€”Teacher: Frozen ModernBERT-large (421M params, 28 layers, d=1024, 16 attention heads).
  • โ€”Student: ModernBERT-base (149M params, 22 layers, d=768, 12 attention heads).
  • โ€”Loss Function: Multi-task joint loss: L = ฮฑ_KD ยท ฯ„ยฒ ยท L_KD(ฯƒ(z_S/ฯ„), ฯƒ(z_T/ฯ„)) + ฮฑ_CE ยท L_CE(z_S, y) with distillation temperature ฯ„ = 2.0, ฮฑ_KD = 0.6, ฮฑ_CE = 0.4.
  • โ€”Optimizer: AdamW (lr=2.5e-5) with Cosine Annealing learning rate schedule.
  • โ€”Hardware: Trained natively on Apple Silicon Metal (MPS).

๐Ÿ’ป Quickstart & Inference

Using the Python Client

python
import laya

# Initialize the distilled 149M model on Apple Silicon MPS or CUDA
agent = laya.Agent("knpatil/laya-browser-agent-base", device="mps")

state = {
    "page": {
        "url": "https://huggingface.co/models",
        "title": "Hugging Face Models",
        "text": "Explore over 1M open-source AI models and datasets."
    },
    "elements": [
        {"index": "1", "role": "textbox", "label": "Search models, datasets, users..."},
        {"index": "2", "role": "link", "label": "Tasks"},
        {"index": "3", "role": "link", "label": "Libraries"}
    ]
}

questions = {
    "action_type": {
        "type": "choice",
        "instructions": "Given the goal 'Search for ModernBERT models', what immediate browser action should be taken?",
        "criteria": {
            "click": "Click a visible link, button, or tab",
            "type": "Enter search text into an input field",
            "scroll": "Scroll down to reveal more content",
            "wait": "Wait for dynamic content to load"
        }
    }
}

prediction = agent.predict(state, questions)
print("Action Decision:", prediction["answers"]["action_type"]["choice"])
print("Confidence:", prediction["answers"]["action_type"]["confidence"])

๐Ÿ“œ Citation & Credits

Developed as part of the BroPilot autonomous browser companion project.