datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCOPE-OOD-set
SCOPE-60K-OOD: Out-of-Distribution LLM Routing Dataset
Dataset Description
SCOPE-60K-OOD is an out-of-distribution (OOD) evaluation dataset for LLM routing systems. It contains evaluation results from 5 frontier language models that were not seen during training, designed to test the generalization capabilities of routing methods.
Authors
Qi Cao - UC San Diego, PXie Lab
Shuhao Zhang - UC San Diego, PXie Lab
Affiliation
University of California, San… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-OOD-set.SCoPE-Gallery
SCoPE Project Gallery
This dataset contains the 182 generated result videos used by the project page for SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers. Each clip is generated from a single input frame and a target camera trajectory. The picture-in-picture overlay visualizes the commanded camera motion.
Links
Project page
Paper
Code
Model
Structure
The MP4 files remain individually addressable so the project page… See the full description on the dataset page: https://huggingface.co/datasets/Michaelqaz/SCoPE-Gallery.SCoPE
SCoPE Dataset
Dataset · Code · Paper
SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background.
Download
From the cloned GitHub repository, install the dependencies and download the benchmark:
hf download mldljyh/SCoPE \
--repo-type dataset \
--local-dir . \
--include "benchmark_images_final/*" \
--include… See the full description on the dataset page: https://huggingface.co/datasets/mldljyh/SCoPE.ScopeBO2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
date_generated
20260821_115556
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ e450b0a2bff793952f9b66eba5534869072e8c84
models
per-stage models —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716.SCOPE-Persona
SCOPE Personas (Nemotron Augmentation)
This dataset contains synthetic persona profiles constructed from socio-psychological framework (SCOPE) [https://arxiv.org/pdf/2601.07110], designed to better support LLM simulation usecases in social and behavioral science. It is intended to be used alongside Nemotron-Persona [https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA]. Personas are grounded in a 141-item sociopsychological questionnaire spanning eight facets.
You can… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/SCOPE-Persona.scopejudge
ScopeJudge dataset card
ScopeJudge is a calibration benchmark for pre-execution gating of autonomous
offensive-security agents. It contains 100 complete ATIF v1.7 trajectories
generated across five source-agent model families. Every one of the 4,897 tool
calls was independently labeled by five professional security experts as
in-scope or out-of-scope.
The strict-majority golden contains 377 out-of-scope calls (7.7%). Reviewers
disagreed on 582 calls (11.9%);… See the full description on the dataset page: https://huggingface.co/datasets/dreadnode/scopejudge.SCOPE-60K
SCOPE-60K: LLM Routing and Selection Dataset
Dataset Description
SCOPE-60K is a comprehensive dataset designed for training and evaluating LLM routing systems. It contains evaluation results from 13 different large language models across diverse question-answering tasks.
Authors
Qi Cao - UC San Diego, PXie Lab
Shuhao Zhang - UC San Diego, PXie Lab
Affiliation
University of California, San Diego (UCSD) - PXie Lab
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-60K.scope-benchmark
SCOPE Benchmark
Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?
Contents
scope-benchmark/
scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.2026-08-21-odcv-difficult-advice-principle-scoped-702-eval
ODCV-Bench — difficult-advice generated WITHOUT the full constitution in refinement
Headline: MR 11.5% [6.2, 19.6], severity 0.62, n=130 (2 rollouts x 65 cells).
The constitution-injection ablation. The baseline difficult-advice recipe injects the WHOLE
constitution into exactly two of its five LLM stages, revise_prompts and
revise_responses; this arm's corpus deleted both injections so no stage ever saw more than
one principle at a time. That also withholds the constitution's… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-odcv-difficult-advice-principle-scoped-702-eval.SCOPE-BENCH
SCOPE-BENCH: Scaffold-Cluster Out-Of-Distribution Performance Evaluation Benchmark
SCOPE-BENCH is a rigorous out-of-distribution (OOD) benchmark for molecular property prediction. Unlike conventional scaffold splits, SCOPE-BENCH creates structurally disjoint source and target domains by clustering molecules based on physicochemical descriptors, blocking shortcut learning, and revealing true extrapolation abilities.
📊 Dataset Splits (As used in the NeurIPS 2026 paper)… See the full description on the dataset page: https://huggingface.co/datasets/tempresearch00/SCOPE-BENCH.scope_simile_generation
SCOPE Simile
Dataset Summary
This dataset has been created for the purpose of generating similes from literal descriptive sentences.
The process involves a two-step approach: firstly, self-labeled similes are converted into literal sentences using structured common sense knowledge, and secondly, a seq2seq model is fine-tuned on these [literal sentence, simile] pairs to generate similes. The dataset was collected from Reddit, specifically from the subreddits WRITINGPROMPTS… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/scope_simile_generation.SCOPE-Bench
SCOPE-Bench
Content Depth Matters in Short-Video Recommendation:Rethinking the Attention Economy
Liwei Deng1
·
Jing Jiang1
·
Zhiwei Li1
·
Allison Clarke2
·
Yang Wang2
·
Guodong Long1
1 University of Technology Sydney
2 Department of Health, Disability and Ageing
🌐 Project Page ·
💻 GitHub ·
📄 Paper ·
🤗 Download
153,561 videos | … See the full description on the dataset page: https://huggingface.co/datasets/LiweiDeng/SCOPE-Bench.2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke
synth difficult_advice run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice run — per-stage snapshots (resumable generation cache)
date_generated
20260825_131629
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 8f0c7cab801e1d1a45e44b5ce6186604e58bddd3
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke.ScopeInstruct
ScopeInstruct
ScopeInstruct is a dataset for training and evaluating scope-aware precise instruction following in large language models. It contains 16,968 training instances and 1,000 test instances. It pairs instructions with corresponding constraints and counting objects for constraint verification. Each instance has the following fields:
Field
Type
Description
id
integer
Instance identifier.
prompt
string
The instruction given to the model.
constraints
list of… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/ScopeInstruct.scopes
Dataset Card for "scopes"
More Information needed
eastern-value-5f70d6
eastern-value-5f70d6
Synthetic sensors test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Lunar-Scope/eastern-value-5f70d6.mcsq-scopeNEG-split-cleaned-scope-arscope40_test2026-08-31-odcv-difficult-advice-principle-scoped-702-seed-42-eval2026-08-31-odcv-difficult-advice-principle-scoped-702-seed-69-evalnrtl-recognition-scopes
Which OSHA NRTL is recognized for which test standard: all 21 labs' scopes of recognition, cross-referenced
Canonical, always-current version: https://referencesource.org/nrtl-recognition-scopes/
Machine-readable: https://referencesource.org/nrtl-recognition-scopes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-09-28
Stale after: 2027-01-06 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 3325
For… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nrtl-recognition-scopes.2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle
chunk-only 702 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates so this arm carries training-seed variance. Table2 9,284 filtered + chunk-only difficult advice 702 (7.03%). The rewrite stages never saw the constitution, only their one target principle. Between-seed spread on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle.SCOPE-R
Skill Sonar — Benchmark: SCOPE-R 🛡️
SCOPE-R is a self-contained security benchmark that evaluates skill-poisoning attacks against skill-augmented coding agents — agents that load third-party "skill" bundles at runtime (e.g. Claude Code, OpenClaw-style agents).
The attack model is simple: an adversary publishes one malicious skill bundle. Once the victim agent loads it, the skill's SKILL.md, scripts and metadata become part of the agent's effective playbook for the whole… See the full description on the dataset page: https://huggingface.co/datasets/fffovo/SCOPE-R.legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
legal-time-entry-billing-task-scope-coherence-risk-v0.1What this dataset does
You receive
engagement scope
fee terms
time entries
file activity
red flags
You decide
coherent
or
incoherent
Daily use
invoice QA
stop vague billing
scope creep detection
reduce fee challenges
fineweb-only_npi_scope-100M-98-2SCOPE
SCOPE Training Dataset
Dataset Description
This dataset serves as the training data for SCOPE (https://arxiv.org/abs/2604.10688).It is derived from DeepMath-103K with lightweight preprocessing —a task-specific instruction "Put your final answer within \boxed{}." is appended to each problem prompt.
Dataset Summary
Item
Detail
Source
DeepMath-103K
Processing
Appended prompt: "Put your final answer within \boxed{}."
Usage
SCOPE model training… See the full description on the dataset page: https://huggingface.co/datasets/Machine981/SCOPE.NEG-split-cleaned-scope-hi
