datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PP
SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring
SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an
operating integrated steel plant. It is designed to evaluate vision-language
models (VLMs) on real-world industrial action recognition, PPE assessment,
and safety-violation detection — under naturally occurring degradation
(dust, glare, steam, low light), at distances and crowdedness levels that
curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.oxford-pets-grpo-think
Oxford-IIIT Pet — GRPO training data (structured reasoning)
GRPO (verl) training data for Oxford-IIIT Pet breed classification with a structured-reasoning prompt: the model emits a scratchpad tagging visible properties (HasProperty), parts (HasA), and setting (AtLocation) before the label. Reward: 0.30 for a well-formed think block, 0.70 for the label match.
Splits: train 2,944 rows, test 3,669 rows.
Schema
column
type
data_source
string
prompt… See the full description on the dataset page: https://huggingface.co/datasets/jucamohedano/oxford-pets-grpo-think.
