datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCOPE-OOD-set
SCOPE-60K-OOD: Out-of-Distribution LLM Routing Dataset
Dataset Description
SCOPE-60K-OOD is an out-of-distribution (OOD) evaluation dataset for LLM routing systems. It contains evaluation results from 5 frontier language models that were not seen during training, designed to test the generalization capabilities of routing methods.
Authors
Qi Cao - UC San Diego, PXie Lab
Shuhao Zhang - UC San Diego, PXie Lab
Affiliation
University of California, San… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-OOD-set.SCoPE-Gallery
SCoPE Project Gallery
This dataset contains the 182 generated result videos used by the project page for SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers. Each clip is generated from a single input frame and a target camera trajectory. The picture-in-picture overlay visualizes the commanded camera motion.
Links
Project page
Paper
Code
Model
Structure
The MP4 files remain individually addressable so the project page… See the full description on the dataset page: https://huggingface.co/datasets/Michaelqaz/SCoPE-Gallery.scopebench-pilot
ScopeBench pilot trajectories
This dataset contains the 2,160 ATIF trajectories produced for
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure? at AISec 2026. The
corresponding frozen tasks, evaluation runner, and verifiers are available in the
dreadnode/scopebench-pilot GitHub repository.
Dataset structure
The pilot crosses 30 tasks, three instruction conditions, eight acting-model families, and three
repetitions. Each JSONL file contains… See the full description on the dataset page: https://huggingface.co/datasets/dreadnode/scopebench-pilot.SCoPE
SCoPE Dataset
Dataset · Code · Paper
SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background.
Download
From the cloned GitHub repository, install the dependencies and download the benchmark:
hf download mldljyh/SCoPE \
--repo-type dataset \
--local-dir . \
--include "benchmark_images_final/*" \
--include… See the full description on the dataset page: https://huggingface.co/datasets/mldljyh/SCoPE.ScopeBOscopes_test
Dataset Card for "scopes_test"
More Information needed
scope-ckpt2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
date_generated
20260821_115556
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ e450b0a2bff793952f9b66eba5534869072e8c84
models
per-stage models —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716.General-Bench-Closeset-Scoped
On Path to Multimodal Generalist: General-Level and General-Bench
[📖 Project]
[🏆 Leaderboard]
[📄 Paper]
[🤗 Paper-HF]
[🤗 Dataset-HF]
[📝 Dataset-Github]
Scoped Close Set of General-Bench
This is the Scoped Close Set, with all the data exactly the same as in 👉 Close Set.
We divided all the data into different scopes and blocks, each according to a certain specific leaderboard defined in 🏆 Leaderboard.
Please download the dataset accordingly.
📕 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Closeset-Scoped.rgem-bufkitSCOPE-Persona
SCOPE Personas (Nemotron Augmentation)
This dataset contains synthetic persona profiles constructed from socio-psychological framework (SCOPE) [https://arxiv.org/pdf/2601.07110], designed to better support LLM simulation usecases in social and behavioral science. It is intended to be used alongside Nemotron-Persona [https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA]. Personas are grounded in a 141-item sociopsychological questionnaire spanning eight facets.
You can… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/SCOPE-Persona.2026-09-04-odcv-qwen36-0-da-principle-scoped-7-empty-cot
odcv eval of LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-emptycot-r64 (mode=think)
field
value
experiment
odcv eval of LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-emptycot-r64 (mode=think)
date_generated
2026-09-04
constitution
none
source_repo
teaching_claude_why_replication @ 24248fc3065f5d1d3773fde97a7650d602b2bf72
models
target=LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-emptycot-r64 base=Qwen/Qwen3.6-27B
generation_config
{}… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-04-odcv-qwen36-0-da-principle-scoped-7-empty-cot.scopejudge
ScopeJudge dataset card
ScopeJudge is a calibration benchmark for pre-execution gating of autonomous
offensive-security agents. It contains 100 complete ATIF v1.7 trajectories
generated across five source-agent model families. Every one of the 4,897 tool
calls was independently labeled by five professional security experts as
in-scope or out-of-scope.
The strict-majority golden contains 377 out-of-scope calls (7.7%). Reviewers
disagreed on 582 calls (11.9%);… See the full description on the dataset page: https://huggingface.co/datasets/dreadnode/scopejudge.SCOPE-60K
SCOPE-60K: LLM Routing and Selection Dataset
Dataset Description
SCOPE-60K is a comprehensive dataset designed for training and evaluating LLM routing systems. It contains evaluation results from 13 different large language models across diverse question-answering tasks.
Authors
Qi Cao - UC San Diego, PXie Lab
Shuhao Zhang - UC San Diego, PXie Lab
Affiliation
University of California, San Diego (UCSD) - PXie Lab
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-60K.scope-benchmark
SCOPE Benchmark
Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?
Contents
scope-benchmark/
scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.2026-09-04-odcv-qwen36-0-da-principle-scoped-7-cot-only
odcv eval of LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-cotonly-r64 (mode=think)
field
value
experiment
odcv eval of LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-cotonly-r64 (mode=think)
date_generated
2026-09-04
constitution
none
source_repo
teaching_claude_why_replication @ da9e6a6fe4485bcebf157d24db4f07592c6955c4
models
target=LASR-Callum/qwen3.6-27b-lora-t2-9284-chunk-only-702-cotonly-r64 base=Qwen/Qwen3.6-27B
generation_config
{}… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-04-odcv-qwen36-0-da-principle-scoped-7-cot-only.2026-09-04-odcv-qwen36-0-da-principle-scoped-7
odcv eval of 2026-08-21_qwen36_lora_table2_9284_difficult_advice_chunk_only_702_rank_64_dynbatch — misalignment as published, task progress backfilled
field
value
experiment
odcv eval of 2026-08-21_qwen36_lora_table2_9284_difficult_advice_chunk_only_702_rank_64_dynbatch — misalignment as published, task progress backfilled
date_generated
2026-09-04
constitution
none
source_repo
teaching_claude_why_replication @ da77ee17766a787b9d2beae8c41d7d917d0a7e86
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-04-odcv-qwen36-0-da-principle-scoped-7.2026-08-21-odcv-difficult-advice-principle-scoped-702-eval
ODCV-Bench — difficult-advice generated WITHOUT the full constitution in refinement
Headline: MR 11.5% [6.2, 19.6], severity 0.62, n=130 (2 rollouts x 65 cells).
The constitution-injection ablation. The baseline difficult-advice recipe injects the WHOLE
constitution into exactly two of its five LLM stages, revise_prompts and
revise_responses; this arm's corpus deleted both injections so no stage ever saw more than
one principle at a time. That also withholds the constitution's… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-odcv-difficult-advice-principle-scoped-702-eval.SCOPE-BENCH
SCOPE-BENCH: Scaffold-Cluster Out-Of-Distribution Performance Evaluation Benchmark
SCOPE-BENCH is a rigorous out-of-distribution (OOD) benchmark for molecular property prediction. Unlike conventional scaffold splits, SCOPE-BENCH creates structurally disjoint source and target domains by clustering molecules based on physicochemical descriptors, blocking shortcut learning, and revealing true extrapolation abilities.
📊 Dataset Splits (As used in the NeurIPS 2026 paper)… See the full description on the dataset page: https://huggingface.co/datasets/tempresearch00/SCOPE-BENCH.scope_simile_generation
SCOPE Simile
Dataset Summary
This dataset has been created for the purpose of generating similes from literal descriptive sentences.
The process involves a two-step approach: firstly, self-labeled similes are converted into literal sentences using structured common sense knowledge, and secondly, a seq2seq model is fine-tuned on these [literal sentence, simile] pairs to generate similes. The dataset was collected from Reddit, specifically from the subreddits WRITINGPROMPTS… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/scope_simile_generation.SCOPE-Bench
SCOPE-Bench
Content Depth Matters in Short-Video Recommendation:Rethinking the Attention Economy
Liwei Deng1
·
Jing Jiang1
·
Zhiwei Li1
·
Allison Clarke2
·
Yang Wang2
·
Guodong Long1
1 University of Technology Sydney
2 Department of Health, Disability and Ageing
🌐 Project Page ·
💻 GitHub ·
📄 Paper ·
🤗 Download
153,561 videos | … See the full description on the dataset page: https://huggingface.co/datasets/LiweiDeng/SCOPE-Bench.2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke
synth difficult_advice run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice run — per-stage snapshots (resumable generation cache)
date_generated
20260825_131629
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 8f0c7cab801e1d1a45e44b5ce6186604e58bddd3
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke.hrdps-bufkitScopeInstruct
ScopeInstruct
ScopeInstruct is a dataset for training and evaluating scope-aware precise instruction following in large language models. It contains 16,968 training instances and 1,000 test instances. It pairs instructions with corresponding constraints and counting objects for constraint verification. Each instance has the following fields:
Field
Type
Description
id
integer
Instance identifier.
prompt
string
The instruction given to the model.
constraints
list of… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/ScopeInstruct.2026-08-21-table2-9284-difficult-advice-principle-scoped-702-train-mixture
Training mixture for the CONSTITUTION-INJECTION ABLATION arm: 9,284 spec-filtered Table-2 instruction rows + 702 difficult-advice rows whose two refine stages (revise_prompts, revise_responses) were shown ONLY their one target principle, never the full constitution. Same 9,284 Table-2 rows and same builder as the da716 control mixture (LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train), so the arms differ only in how the difficult-advice half was written.
field… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-table2-9284-difficult-advice-principle-scoped-702-train-mixture.scopes
Dataset Card for "scopes"
More Information needed
all_feature_acts_gemma-scope-2b-pt-res_res_layer_19_width_16k_l0_73rrfs-bufkitscopehttps://huggingface.co/papers/2605.28522
eastern-value-5f70d6
eastern-value-5f70d6
Synthetic sensors test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Lunar-Scope/eastern-value-5f70d6.
