datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agents-last-exam-reference
Agents Last Exam — Reference (Ground-Truth) Data
⚠️ Gated dataset. This repo contains the ground-truth / reference outputs
used to score the Agents Last Exam (ALE) benchmark. Access requires login,
agreement to the terms on the access-request form, and manual approval.
Note (06/16/26): This repository was accidentally deleted and has been recreated. The
previous list of approved requesters could not be restored, so even if you
were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.tadabur-align-references
tadabur-align-references
Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR).
What this is
For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.terrain-referenced-3d-glacier-mapping-product
Terrain-referenced Glacier Mapping Product
This repository provides the terrain-referenced glacier-area mapping product generated for the manuscript. The product is openly available through Hugging Face with DOI: 10.57967/hf/9900.
The archive contains regional mapping outputs, oblique terrain-visualization products, metadata files, and tabular glacier-area summaries. Glacier masks generated by Prithvi-SDT are linked with Copernicus DEM terrain information and RGI 7.0 glacier… See the full description on the dataset page: https://huggingface.co/datasets/yyhw/terrain-referenced-3d-glacier-mapping-product.multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.watercolour-reference-pool
Watercolour reference pool
The reference paintings that define the reward in the watercolour RL environment: an
agent writes a p5.brush sketch, the sketch is
rendered, and a vision judge compares the render against paintings sampled from this pool.
What the pool contains is the reward function. Replace it and you have changed what
the environment rewards, without touching a line of code.
178 paintings in two tiers, each with the JavaScript source that produced it.
tier… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-reference-pool.references
GEM References
What is it?
This repository contains all the reference datasets that are used for running evaluation on the GEM benchmark. Some of these datasets were originally hosted as a GitHub release on the GEM-metrics repository, but have been migrated to the Hugging Face Hub.
Converting datasets to JSON
We provide a convert_dataset_to_json.py conversion script that converts the datasets in the GEM organisation to the JSON format expected by the… See the full description on the dataset page: https://huggingface.co/datasets/GEM/references.moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.xet-spec-reference-filesThe files in this repository are intended to provide a reference to content processed using the xet protocol, relative to the original file: Electric_Vehicle_Population_Data_20250917.csv
The original file was exported from https://data.wa.gov/Transportation/Electric-Vehicle-Population-Data/f6w7-q2d2/about_data on September 16, 2025.
The contents are as described:
Electric_Vehicle_Population_Data_20250917.csv - the original file
Electric_Vehicle_Population_Data_20250917.csv.chunks - a… See the full description on the dataset page: https://huggingface.co/datasets/xet-team/xet-spec-reference-files.low-high-reference
Reference Directory
git clone https://github.com/PKU-YuanGroup/Helios.git
git clone https://github.com/NVlabs/LongLive.git
git clone git@github.com:bingreeky/MemGen.git
这个目录用于存放项目设计、实现和训练过程中会反复参考的外部资料。它不是运行时必须的源码目录,而是研究与开发参考层。
目录定位
reference/ 主要存放以下几类内容:
论文 PDF
论文配套笔记
方案草稿
外部开源项目的结构化阅读记录
和当前项目直接相关的训练/规划/critic/memory 参考材料
它的作用不是“被 import”,而是帮助回答这些问题:
当前系统应该如何拆成 planner / edit / critic / memory
哪些训练阶段适合先做监督、后做偏好、再做 RL / GRPO
哪些工作可以作为 skill / memory / reflection… See the full description on the dataset page: https://huggingface.co/datasets/Ouzhang/low-high-reference.sai-osworld-v21-reference-runs
Sai on OSWorld v2.1 — reference-VM runs
Every clean run of Sai (agent model anthropic/claude-opus-5, thinking effort max) on the 108 tasks of the
osworld-v2.1 release, on a cocoon microVM built from the official v2.1 reference image (Ubuntu 22.04.3).
134 runs across 108 tasks: some tasks ran more than once, and every clean run is here.
Results
108/108 tasks have at least one clean run.
Headline (mean over tasks of each task's mean clean score): 0.7731.
First… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v21-reference-runs.tldr-with-sft-referenceMedical-Data-Referenceasset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.GWASLab-Reference
GWASLab reference datasets
Processed genomic reference files used by GWASLab (download_ref / gwaslab download ref). This dataset replaces the previous Dropbox hosting for GWASLab-processed panels. Official dbSNP VCFs, UCSC FASTA, Ensembl/RefSeq GTF, and liftOver chains stay at their original hosts.
Package catalog: reference.json. Checksums for every file in this repo are in md5sum.txt.
Download with GWASLab
import gwaslab as gl
gl.download_ref("1kg_eas_hg19")… See the full description on the dataset page: https://huggingface.co/datasets/Cloufield/GWASLab-Reference.asr-reference-set-eval-temp
Temporary ASR evaluation audio
Temporary public audio files used for hosted ASR evaluation.
moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.kimi-k3-full-mxfp4-kld-reference-32x2048
Kimi K3 full-MXFP4 KLD reference logits
This dataset contains the canonical full-vocabulary reference logits for
quantization comparisons of Kimi K3. The source is the original full MXFP4
checkpoint served as W4A16 on TP16 with vLLM dev/gg-k3, SparkInfer, and
InstantTensor.
Contents
32 independent 2048-token windows
65,504 scored next-token positions (32 * 2047)
vocabulary size 163,840
one [2047, 163840] F32 safetensors tensor per window
tensor key: logits
total… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-full-mxfp4-kld-reference-32x2048.ssim-reference-videosseqqa-reward-reference
SEQQA reward reference
Regression fixture for trl.internal.seqqa.reward, used by tests/internal/test_seqqa_rewards.py.
39 hand-checked answers across 20 SEQQA validator types: for each the
answer string, the ideal reference answer, and whether seqqa_accuracy_reward is expected to score it
1.0 (expected = true) or 0.0. validator_params_json is the serialised validator payload, and
source_id / source_revision point at the generated question this row was pinned from.
Values are… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/seqqa-reward-reference.bullinger-references-topicsspeedpainting-reference-expansion-v1
Speedpainting reference expansion
This repository archives licensed reference photographs, explicit bounded-brush-v2 programs, canonical 600×600 node renders, and independent assistant reviews for the reference-to-painting project.
The September 14 pilot contains 16 new reference photographs. After one geometry-repair round, 15 teacher targets were admitted for coarse structure supervision; the orchard remains held. These are manually composed Astra teacher programs, not Qwen… See the full description on the dataset page: https://huggingface.co/datasets/CK0607/speedpainting-reference-expansion-v1.test_referencesreference-logitsThis is just a temp scratchpad for sharing data.
peg_rand_05_01_cam_reference_cam0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_cam_reference_cam0.pill-reference-imagesFrontierChallenge-reference
FrontierChallenge reference data
FrontierChallenge reference data provides 97 authenticated, encrypted
verifier archives.
Path
Contents
tasks/<task-id>/verifier.fcref
encrypted tests/: grader, rubric, fixtures, validation code, and reference outputs
manifest.jsonl
archive paths, sizes, and SHA-256 commitments
source_registry.json
release binding shared with GitHub and the solve dataset
tools/
integrity checker and standalone unsealer
The archive password is… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge-reference.6k-diverse-reference-voices
6k Diverse Reference Voices
6,064 permissively licensed reference voices for casting expressive voice-acting generations.
All voices in this collection are permissively usable: they were either synthetically created or
extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0.
Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0),
re-published under LAION with clarified metadata documentation. If you use this dataset,
please attribute… See the full description on the dataset page: https://huggingface.co/datasets/laion/6k-diverse-reference-voices.state-civil-statute-of-limitations
Civil statute of limitations by state and type of claim
Canonical, always-current version: https://referencesource.org/state-civil-statute-of-limitations/
Machine-readable: https://referencesource.org/state-civil-statute-of-limitations/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-25
Stale after: 2027-08-25 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 238
How long do you have to sue? Every… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-civil-statute-of-limitations.iso-standard-supersessions
Withdrawn and superseded ISO standards
Canonical, always-current version: https://referencesource.org/iso-standard-supersessions/
Machine-readable: https://referencesource.org/iso-standard-supersessions/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-05
Stale after: 2027-03-13 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 24640
Every deliverable in ISO's own open-data register that has been… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/iso-standard-supersessions.TTS_Reference
