Team Ai
Modelpublic

Lexmount/WebJev-35B-A3B

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
1likes34downloads
Model Card

WebJev-35B-A3B

WebJev-35B-A3B is a decision model built to drive browser agents on real websites.

At each step of a web task, WebJev reads the agent's view of the page and makes the step's two decisions:

  • —which action to take next;
  • —which of up to 255 on-page elements to take it on.

The page view is its URL, its visible text, its actionable elements and the recent actions.

WebJev is fine-tuned from Qwen3.5-35B-A3B-Base as a discriminative decision model: instead of generating text, it learns to score the options of a question. Its training combines two sources:

  • —tens of thousands of execution-verified decisions that browser agents made on live websites;
  • —a broad mixture of typed decisions: classification, routing, tool choice, policy and evidence checking, and knowledge-intensive multiple choice.

The result is precise element grounding on long, cluttered pages, reliable next-action choices and strong general structured decisions. Each decision takes a single forward pass and returns a full probability distribution over the options.

**Code** · **Dataset** · **Training**

[image]

Highlights

  • —Next-action prediction and element grounding. WebJev decides directly from the agent's page state:
  • —the action: click, type, select, scroll, wait, submit, dismiss, go back, finish or give up;
  • —the element to act on, among up to 255 candidates on long real-world pages.

It learns both from execution-verified decisions on live websites (Lexmount/WebJev).

  • —Stronger agents on the live web. In the same browser agent, with the same tasks and budget, WebJev completes 38.5% of 125 live-website tasks. jev-1.13 completes 16.7%, so WebJev solves 2.3× as many. It leads on all three task collections, with four times the success rate on WebGym.
  • —General structured decisions. WebJev leads jev-1.13 on JevBench public (87.9 against 85.7) and on multi-class classification (86.3 against 83.3). It is ahead on Nimble evidence checking and customer-ticket triage too.
  • —Fast and exact. Each decision is one forward pass with no decoding: a web-page decision takes about a third of a second on one A100. The answer is always one of the listed options.

Model overview

WebJev scores the options of each question about a state. It never generates free text, so there is no output parsing and no answer outside the options you list.

Developed byLexmount
Model typedecision model (option scoring at an answer position), sparse mixture of experts
Base modelQwen/Qwen3.5-35B-A3B-Base
Parameters34.66B in total, about 3B active per token
WeightsBF16, 15 safetensors shards, 69.3 GB
Contexttrained on inputs of up to 16,384 tokens
Options per question2–255
Inferencetransformers, vLLM
LanguagesEnglish instructions; English and Chinese web pages
Codegithub.com/lexmount/WebJev
Training dataLexmount/WebJev
LicenseApache-2.0 (see License)

[image]

Evaluation

All numbers are measured with the released BF16 weights served by vLLM, at temperature 1.0. The jev-1.13 numbers were measured through its official API on the same inputs.

End-to-end web tasks

WebJev serves as the decision component of the same browser agent on 125 real-website tasks: 75 from Online-Mind2Web, 26 from WebGym and 24 from WebVoyager. Each run has a budget of 900 seconds and 60 actions. A deterministic grader checks the final page state and answer.

Success rate = solved ÷ evaluable tasks; tasks lost to browser infrastructure or grader errors are excluded.

Task setjev-1.13**WebJev-35B-A3B**
All tasks16.67% (20/120)38.52% (47/122)
Online-Mind2Web16.90% (12/71)38.36% (28/73)
WebGym8.00% (2/25)32.00% (8/25)
WebVoyager25.00% (6/24)45.83% (11/24)

General structured decisions

Eight benchmarks of structured decisions, beyond the web. Each item gives a state and a set of candidate answers, and the model's choice is correct when it equals the reference label. The table reports accuracy.

Benchmark (items)What it measuresjev-1.13**WebJev-35B-A3B**
JevBench public (231)general structured decisions (intent, extraction, tool choice, policy); the hard tier has long policies, multi-hop, temporal and numeric reasoning, and trap items85.7187.88
Multi-class decisions, dev (1,468)news topic, review sentiment and stars, 77-way banking intent, question and entity type, yes/no reading comprehension, entailment, policy rules83.3186.31
Customer-ticket triage (873)routing queue, anger and priority of support tickets (partly Korean)74.9176.29
Nimble held-out (324)fine-grained evidence checking with minimal pairs: one fact changes and the answer flips92.5992.90
SemIf external (252)claim verification: supported, refuted or not enough evidence98.4198.02
Cross-task transfer, dev (764)MMLU, Emotion, TweetEval, QNLI, PAWS, SciQ and programmatic policy-rule questions85.2185.08
Typed business decisions, test (2,000)agent-trajectory stop and escalation, customer requests, invoice approval, security-alert severity; teacher labels74.0573.90
MMLU-Pro, 10 options (1,000)college-level knowledge and reasoning across subjects83.4069.40
Average (equal weights)84.7083.72

Speed

Measured on one NVIDIA A100 80GB with vLLM and BF16. Requests are sent one at a time; latency covers all questions about one state.

Workloadp50p95
short state, one question (claim verification)72 ms84 ms
policy or knowledge question (JevBench, MMLU-Pro)145 ms150–300 ms
support ticket, three questions228 ms309 ms
web page state, one decision337 ms755 ms

How to use

With transformers

The model needs a transformers version with Qwen3.5 MoE support (qwen3_5_moe_text, 5.15 or newer) and one GPU with at least 80 GB of memory.

python
import json, torch
from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "Lexmount/WebJev-35B-A3B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda").eval()

state = {"page": {"url": "https://www.spanishdict.com/", "title": "SpanishDictionary.com", "text": "…"},
         "elements": [{"index": "9", "role": "combobox", "label": "Translate Spanish or English", "value": "spring"},
                      {"index": "14", "role": "option", "label": "spring"}],
         "recent_actions": [{"action": "Translate Spanish or English", "kind": "fill", "text": "spring"}]}
options = ["CLICK", "TYPE_TEXT", "SCROLL_DOWN", "PRESS_ENTER", "DONE", "BLOCKED"]
prompt = ("Context:\n" + json.dumps(state, ensure_ascii=False) +
          "\n\nQuestion: Search SpanishDict for 'spring'. Which operation should the agent perform next?\nOptions:" +
          "".join(f"\n({chr(65 + i)}) {o}" for i, o in enumerate(options)) + "\nAnswer: (")

inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
    logits = model(**inputs).logits[0, -1]
label_ids = [tok.encode(chr(65 + i), add_special_tokens=False)[0] for i in range(len(options))]
probs = torch.softmax(logits[label_ids].float() / 1.0, -1)       # temperature from decider_config.json
print(dict(zip(options, probs.tolist())))

For questions with more than 10 options, labels continue as single tokens (K … Z, then AA, AB, …). Build those prompts with decider/prompt.py from this repository, which renders every label as one token.

With vLLM

Start a server that returns the logits of the allowed tokens:

bash
vllm serve Lexmount/WebJev-35B-A3B --dtype bfloat16 --max-model-len 34816 \
    --logprobs-mode processed_logits --max-logprobs 256

Then score a prompt: one generated token, restricted to the option labels.

python
import math, requests

ids = tok(prompt, add_special_tokens=False)["input_ids"]          # the prompt above, ending with "Answer: ("
out = requests.post("http://127.0.0.1:8000/v1/completions", json={
    "model": "Lexmount/WebJev-35B-A3B", "prompt": ids, "max_tokens": 1, "temperature": 1.0,
    "logprobs": len(options), "allowed_token_ids": label_ids, "return_tokens_as_token_ids": True}).json()
top = out["choices"][0]["logprobs"]["top_logprobs"][0]           # {"token_id:<id>": logit}
logits = [top[f"token_id:{i}"] for i in label_ids]
z = [math.exp(x - max(logits)) for x in logits]
print(dict(zip(options, [x / sum(z) for x in z])))

How it works

The input is Context: {state}, followed by the question, the lettered options (A) … (B) … and the answer position Answer: (. At that position, the model's hidden state is projected onto the LM-head rows of the option labels only and normalized with a softmax over the valid labels. The labels are never generated.

  • —Several questions about one state are scored as separate prompts that share the state as a prefix.
  • —Option order is part of the input. The model was trained with shuffled options, so it conditions on the candidates' content rather than their position.
  • —Inference settings are stored in decider_config.json: temperature 1.0, up to 255 options per question, state before question.

Model architecture

ArchitectureQwen3_5MoeForCausalLM (text model)
Layers40: 10 full-attention layers and 30 Gated DeltaNet linear-attention layers (one full-attention layer every four)
Hidden size2,048
Full attention16 query heads, 2 key-value heads, head dimension 256, gated output
Linear attention16 key heads, 32 value heads, head dimension 128
Experts256 routed experts per layer (8 active per token) and 1 shared expert, expert width 512
Vocabulary248,320 tokens
Parameters34,660,610,688 in total, about 3B active per token

Training

Intended uses

Intended uses.

  • —The decision component of browser agents: next-action prediction and element grounding.
  • —Typed decisions in software pipelines: routing, classification, extraction choices, policy and evidence checks. Each is a question with an explicit option list.

Out of scope.

  • —Free-form generation, chat and open-ended question answering.
  • —Decisions whose options are not listed in the input.

Usage notes.

  • —Knowledge-intensive decisions. WebJev decides from the state, so include the relevant facts in it.
  • —Confidence. Probabilities use temperature 1.0 and rank options reliably. To act on confidence thresholds, check calibration on your own labels.
  • —Input and hardware. Inputs of up to 16,384 tokens. One GPU with 80 GB of memory runs the BF16 weights.
  • —Live websites change over time, so end-to-end results on live tasks vary between runs.

License

WebJev-35B-A3B is released under the Apache License 2.0. You may download, use, fine-tune and redistribute the weights, including for commercial use.

  • —The model is a fine-tuned derivative of Qwen3.5-35B-A3B-Base, released under the Apache License 2.0.
  • —The prompt builder in decider/ derives from the open-source Decider package, under the same license.

Citation

bibtex
@misc{lexmount2026webjev,
  title        = {WebJev-35B-A3B: A One-Pass Decision Model for Web Agents},
  author       = {{Lexmount}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Lexmount/WebJev-35B-A3B}},
  url          = {https://github.com/lexmount/WebJev}
}

Contact

Lexmount, via the Lexmount organization on Hugging Face.