datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-task-recursive-task-synthesis
Recursive-Task-Synthesis for tmax
Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments.
All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.Recursive-Task-Synthesis
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.Recursive-Task-Synthesis
Recursive Task Synthesis
Tasks without completed platform artifacts or with unresolved VM validation
failures are temporarily excluded. exclusions.json records the exact IDs,
reasons, build IDs where available, and evidence dates/runs. Exclusions affect
both metadata rows and complete TAR task packages. Runtime failures are not
image-build failures or proof of incorrect gold solutions. This filter does
not establish that every retained task passes gold validation.
Restore a task… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Recursive-Task-Synthesis.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.Recursive-Task-Synthesis-Trajectories
Recursive Task Synthesis Trajectories
This dataset contains 327,189 completed agent trajectories collected on
recursively synthesized command-line tasks. Public identifiers are opaque and
stable.
The trajectory JSON retains messages, actions, observations, and token counts.
Token-level log-probability arrays and duplicated debug/session captures are
excluded from the public packages.
metadata/trajectories.parquet: searchable trajectory metadata.
metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.synthesis
NuBerea/synthesis
A cross-corpus synthesis layer for the study of early Jewish and Christian literature.
Each config joins pericope-level text units from one corpus — the canonical Bible
(Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag
Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted
from those units and with links into a shared concept vocabulary. The result is a set
of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.v1
Knowledge in Visual Synthesis
This dataset contains prompt–image examples for evaluating and studying
knowledge-intensive visual synthesis. Samples are organized by contributor as
dataset subsets (configs), with each upload version exposed as a split.
Dataset structure
Subset
Splits
byx
v1, v2
yuner
v1, v2, v3
zanyi
v1, v2, v3
jiayu
v1, v2, v3
sherry
v1, v2
yujunz
v1
The byx/v1 split contains 140 unique prompts and 300 generated images.… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.Recursive-Task-Synthesis-Quality-1K
Recursive Task Synthesis Quality 1K
This dataset contains 1,000 quality-selected, validated command-line task
instances. It is a curated subset of the
Recursive Task Synthesis dataset.
Public task and group identifiers are opaque and stable across both datasets.
Selection
The subset was selected from 37,484 validated tasks using structural and safety
checks, two-pass semantic review, strict gates for instruction clarity,
instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.sd_asr_synthesis_dataTowards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modelcommon_voice_17_0_romanian_speech_synthesiscommon_voice_16_1_romanian_speech_synthesissd_asr_synthesis_data_v0_less_silencecommon_voice_romanian_speech_synthesissynthetic_program_synthesis_python_1Mconcepticon-wiktionary-synthesis
Concepticon-Wiktionary Synthesis: Dataset for Tagging the EFEO-CNRS-SOAS Lexicon
CHAPTER I. INTRODUCTION CHAPTER II. FIELDWORK AND COLLECTION STRATEGY CHAPTER III. TECHNICAL METHODOLOGY A. Dataset Synthesis B. Dataset Structure REFERENCES
Abstract
Abstract
A synthesised, multilingual conceptual alignment dataset is presented to facilitate a shared task aimed at the semantic tagging of the EFEO-CNRS-SOAS lexicon with abstract… See the full description on the dataset page: https://huggingface.co/datasets/DTLR/concepticon-wiktionary-synthesis.Force-Controlled-Robotic-Mechanochemical-Synthesiscosmos-trajectory-synthesis
COSMOS Synthetic Traffic Trajectory Dataset
Related releases: Earlier 1042-scene export (incl. normal split) · Real COSMOS trajectories
Description
Synthetic multi-agent traffic trajectories at a fixed urban intersection, generated by the COSMOS pipeline
with Protocol V2 LLM backends and Tier-1 quality gates (geometry + kinematics).
Each scene is created from a natural language prompt and processed through:
Scene Planner → GMM Sampling → Waypoint Filter → Event… See the full description on the dataset page: https://huggingface.co/datasets/JojoZhu/cosmos-trajectory-synthesis.codenib-synthesis
CodeMiner Synthesis
Status: In active development. This dataset is an early work-in-progress.
Both the set of instances and the per-instance query catalog are growing, and
the schema may evolve. Counts shown below describe the current snapshot
only — they are not a final target.
A growing collection of LLM-synthesized natural-language code-search
evaluation queries, each grounded on a real code symbol from a SWE-bench instance.
Design discussion and progress tracking:… See the full description on the dataset page: https://huggingface.co/datasets/sysevol-ai/codenib-synthesis.Heterogenous_Synthesis_Benchmark
Heterogenous_Synthesis_Benchmark
This repository presents a diverse tabular data generation benchmark. We invite you to refer to our paper on arxiv to explore the mechanism behind our data diversity, which we called Distribution-Guided-Rule (DGR). Within this benchmark, you can experience how diverse preference data coverage combined with customized generation enhances post-training performance.
Additionally, Heterogenous_Synthesis_Benchmark includes a comprehensive toolkit for… See the full description on the dataset page: https://huggingface.co/datasets/CurryOvO/Heterogenous_Synthesis_Benchmark.Force-Controlled-Robotic-Mechanochemical-Synthesisvietquill-qcpg-100k-synthesis-question2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus
Description
48kHz, 24bit 품질의 영어 음성 데이터셋으로, 전문 녹음 스튜디오에서 전문 성우 2명(남성 1명, 여성 1명)의 음성을 수집했습니다. 주어진 주제에 대한 즉흥 발화, 다단계 감정, 단일 감정 및 준언어적 특성(Paralinguistic Features) 등 다양한 음성 콘텐츠를 포함합니다.
텍스트, 감정 및 준언어적 특성에 대한 어노테이션을 제공하며, 음성 합성(Speech Synthesis) 등의 음성 AI 모델 개발 및 학습에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/tts/1540?source=hf.kr
Specifications
Format
48kHz, 24bit, 비압축 WAV, 모노 채널
Recording Environment
전문 녹음 스튜디오… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus.Recursive-Task-Synthesis-Copy
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Recursive-Task-Synthesis-Copy.soict_train_synthesis_dataset
Dataset Card for "soict_train_synthesis_dataset"
More Information needed
literary-synthesis
Literary Synthesis
This dataset repurposes the original agentlans/literary-reasoning
data by reformatting it as creative writing prompts paired with literary-style outputs.
Writing style attributes were put in random order, with prompts randomly either prepended or appended.
The output text has been cleaned to make it suitable for creative writing and literary generation tasks.
The rows were sorted by increasing reading difficulty for curriculum learning.
synthesis_data_v1
Dataset Card for "synthesis_data_v1"
More Information needed
