datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
web-page-quality-edu-strict-blind-5196
Web Page Quality EDU Strict-Blind 5196
This dataset contains 5,196 English web pages sampled from
TeraflopAI/web-page-quality-labels
at revision aee0ebeff5bf09e9bd6881b9da581ec6ee77db5c, together with a new
strict-blind educational-quality judgment.
Sampling
The sample uses seed 20260927 and is stratified by the source edu_score:
Source score
Rows
0
1,000
1
1,000
2
1,000
3
1,000
4
1,000
5
196 (all available rows)
Strict-Blind… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/web-page-quality-edu-strict-blind-5196.WEBPRMBENCH
WebPRMBench
The first comprehensive evaluation benchmark for Web Process Reward Models
Published at ICLR 2026
Paper | Code | Website | Collection | Demo
Overview
WebPRMBench is the first comprehensive evaluation benchmark dedicated to Web Process Reward Models (WebPRMs). It evaluates how well a reward model can judge the quality of web agent actions during long-horizon web navigation. Each instance presents a web state (page context, trajectory history, user… See the full description on the dataset page: https://huggingface.co/datasets/ZYao720/WEBPRMBENCH.web-page-quality-labelsweb_pro_backupweb-page-quality-labels-fullweb_pro_plusplus_decont_backupwebpagesankaku_curated_webp-4Mpixel_new
