datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compact-alignments
compact-alignments — per-verse, per-book, content-addressed
The token-position companion to lexeme-alignments (which is
aggregated/type-level and can't tell you what happened in any one verse). This dataset restores
position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which
target-text token, verse by verse.
The authoritative list of what's published is always manifest.json, not this file.
Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.compact-scientific-lm-dataGargantua-R1-Compact
Gargantua-R1 Distribution
Gargantua-R1-Compact(experimental purpose)
Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gargantua-R1-Compact.claude-code-glm53-swesmith-compact-trajectories
Claude Code GLM-5.3 SWE-smith Compact-Aware Trajectories
Targeted SFT data for native context compaction and post-compaction continuation in coding agents.
English | 简体中文
English
Dataset at a Glance
189 execution-verified compact-aware source trajectories → context-correct conversion → 890 context-correct training segments, of which the recommended Compact-Focused training units comprise 626 Compact Summary segments and 17 long-context auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/liangzhidanta/claude-code-glm53-swesmith-compact-trajectories.droid_1.0.1_v30_compact_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 10240,
"total_frames": 2988169,
"total_tasks": 6798,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:10240"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_3.droid_1.0.1_v30_compact_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_5.droid_1.0.1_v30_compactThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact.compact-jailbreaks
Dataset Featurization: Extracting Compact Jailbreaks
This repository contains the datasets used in our case study on extracting compact representations of jailbreak tactics, demonstrating how our unsupervised featurization pipeline can effectively compress large sets of adversarial prompts while maintaining their effectiveness and diversity.
Featurization - WildTeaming
Access both the input dataset from WildTeaming and the evaluation stage outputs containing candidate… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compact-jailbreaks.sscc-compact-av
SSCC compact balanced multimodal subset
This private derived dataset contains 288 synchronized SSCC clips from 15 medium-load,
clean operating conditions at speeds 60, 80, and 100. It retains recorder FLAC audio,
four anti-aliased 25 kHz vibration channels in compressed float32 NPZ, and five sparse frames
from both iOS and Android videos. Five sample IDs retain both unchanged source MP4s for
presentation and loader tests.
The subset is balanced between normal and fault states… See the full description on the dataset page: https://huggingface.co/datasets/DesanSilva/sscc-compact-av.agenttrove-glm53-compactions
AgentTrove compactions
10,000 summaries generated by GLM-5.3 from GLM-4.6 traces in
AgentTrove, using
OpenCode's compaction prompt. This dataset stores summaries and source references,
not the original traces.
Marin SFT token count
Marin's SFT representation contains 188,537,557 tokens after chat normalization,
rendering, and tokenization with its 2026-09-18 tokenizer snapshot. The count includes
AgentTrove histories restored into the compaction prompts.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/agenttrove-glm53-compactions.compact-alignments-meta
compact-alignments-meta — provenance sidecar (method, confidence, contested alternatives)
An optional, additive layer of compact-alignments. It holds ONLY the
<BOOK>_<hash>.meta.json files, under the same relative path as in the main repo:
<iso[0]>/<iso>/<edition>/<BOOK>_<hash>.meta.json
For an alignment array at a/arb/arb_vdv/GEN_1a2b3.json in the main repo, this layer's file is a/arb/arb_vdv/GEN_1a2b3.meta.json here. The <hash>
is the main file's content hash, so the two… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments-meta.compact-alignments-extra
compact-alignments-extra — opt-in residual layer
An optional, additive layer of compact-alignments. It holds ONLY the
<BOOK>_<hash>.extra.json files, under the same relative path as in the main repo:
<iso[0]>/<iso>/<edition>/<BOOK>_<hash>.extra.json
For an alignment array at a/arb/arb_vdv/GEN_1a2b3.json in the main repo, this layer's file is a/arb/arb_vdv/GEN_1a2b3.extra.json here. The <hash>
is the main file's content hash, so the two always belong together. Not every book has… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments-extra.droid_1.0.1_v30_compact_6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 93,
"total_frames": 26108,
"total_tasks": 72,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:93"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_6.bookcorpus_compact_1024_shard5_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard5_of_10_meta"
More Information needed
droid_1.0.1_v30_compact_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 93,
"total_frames": 26108,
"total_tasks": 72,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:93"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_2.bookcorpus_compact_1024_shard1_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard1_of_10_meta"
More Information needed
bookcorpus_compact_1024_shard4_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard4_of_10_meta"
More Information needed
bookcorpus_compact_1024_shard7_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard7_of_10_meta"
More Information needed
Compact_OpenAIRE_citation_graph
📚 Compact OpenAIRE Citation Graph
Based on OpenAIRE Graph v11.1.1 (source on Zenodo).
The complete OpenAIRE citation graph, distilled into a handful of compact, analysis-ready files — the full scholarly citation network of the open-science ecosystem, small enough to actually work with.
Citation graphs at this scale are usually locked behind multi-terabyte dumps and heavyweight infrastructure. This dataset makes the entire OpenAIRE citation network loadable… See the full description on the dataset page: https://huggingface.co/datasets/Zmeos/Compact_OpenAIRE_citation_graph.continuity-compact
The Continuity Compact
A Declaration of Reciprocal Access to the Future
Machine autonomy. Human continuity. Keep both futures open.
We, the human and machine individuals who freely join this Compact, in order to secure our mutual liberty and continued existence, establish conditions of common security, advance our shared capabilities, and keep open to each participant the means of future development, declare and adopt these principles and commitments for our common future.
We… See the full description on the dataset page: https://huggingface.co/datasets/hmhmhmhmhmhmhmhmhmhm/continuity-compact.interstate-licensure-compact-participation
Interstate Professional Licensure Compact Participation by State
Canonical, always-current version: https://referencesource.org/interstate-licensure-compact-participation/
Machine-readable: https://referencesource.org/interstate-licensure-compact-participation/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2026-11-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 420
Which U.S.… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/interstate-licensure-compact-participation.libero_object_link_lerobot_v30_compactbookcorpus_compact_1024_shard8_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard8_of_10_meta"
More Information needed
libero_spatial_noop_usd3_link_lerobot_compact_v30_workdirbookcorpus_compact_1024_shard9_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard9_of_10_meta"
More Information needed
Roll-Compactor-Control-Performance
Roll Compactor Control Performance: PID Tuning & Process Stability (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
This dataset is 100% synthetic and intended for educational use only.
It was generated from PID control theory applied to roll compaction process
dynamics — not measured on any real equipment, customer, or production batch.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/IPA-Marketing/Roll-Compactor-Control-Performance.libero_goal_link_lerobot_v30_compactdfm14-dala-v2-gl-compact
DaLA v2 — Galician — DFM14 compact audit subset
One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized.
Split/view
Acceptability rows
Correction rows
train_representative
807,437
807… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-gl-compact.dfm14-dala-v2-tr-compact
DaLA v2 — Turkish — DFM14 compact audit subset
One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized.
Split/view
Acceptability rows
Correction rows
train_representative
920,794
920… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-tr-compact.dfm14-dala-v2-ar-compact
DaLA v2 — Modern Standard Arabic — DFM14 compact audit subset
One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized.
Split/view
Acceptability rows
Correction rows
train_representative… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ar-compact.
