datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-task-recursive-task-synthesis
Apptainer pool for hamishivi/agent-task-recursive-task-synthesis
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image.
Apptainer images
The pool currently contains 29,501 / 29… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-recursive-task-synthesis.Recursive-Task-Synthesis
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.agent-task-recursive-task-synthesis
Recursive-Task-Synthesis for tmax
Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments.
All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.LongWorld-Synthesis-Workspace
LongWorld synthesis workspace
Machine-migration bundle for continuing synthesis. Train-ready SFT rows
are in Xnhyacinth/LongWorld-Real-Workflows. Probe signing keys and credentials are excluded from this public dataset.
Verification keys must be provisioned separately through the owning workspace.
Layout
Path
What it is
source_inventory/
HMAC-signed source records plus retained source bytes (arXiv tarballs, filings). Cannot be rebuilt without the same… See the full description on the dataset page: https://huggingface.co/datasets/Xnhyacinth/LongWorld-Synthesis-Workspace.Recursive-Task-Synthesis
Recursive Task Synthesis
Tasks without completed platform artifacts or with unresolved VM validation
failures are temporarily excluded. exclusions.json records the exact IDs,
reasons, build IDs where available, and evidence dates/runs. Exclusions affect
both metadata rows and complete TAR task packages. Runtime failures are not
image-build failures or proof of incorrect gold solutions. This filter does
not establish that every retained task passes gold validation.
Restore a task… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Recursive-Task-Synthesis.Recursive-Task-Synthesis-Trajectories
Recursive Task Synthesis Trajectories
This dataset contains 327,189 completed agent trajectories collected on
recursively synthesized command-line tasks. Public identifiers are opaque and
stable.
The trajectory JSON retains messages, actions, observations, and token counts.
Token-level log-probability arrays and duplicated debug/session captures are
excluded from the public packages.
metadata/trajectories.parquet: searchable trajectory metadata.
metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.Recursive-Task-Synthesis-Quality-1K
Recursive Task Synthesis Quality 1K
This dataset contains 1,000 quality-selected, validated command-line task
instances. It is a curated subset of the
Recursive Task Synthesis dataset.
Public task and group identifiers are opaque and stable across both datasets.
Selection
The subset was selected from 37,484 validated tasks using structural and safety
checks, two-pass semantic review, strict gates for instruction clarity,
instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.synthesis
NuBerea/synthesis
A cross-corpus synthesis layer for the study of early Jewish and Christian literature.
Each config joins pericope-level text units from one corpus — the canonical Bible
(Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag
Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted
from those units and with links into a shared concept vocabulary. The result is a set
of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.v1
Knowledge in Visual Synthesis
This dataset contains prompt–image examples for evaluating and studying
knowledge-intensive visual synthesis. Samples are organized by contributor as
dataset subsets (configs), with each upload version exposed as a split.
Dataset structure
Subset
Splits
byx
v1, v2
yuner
v1, v2
zanyi
v1, v2, v3
jiayu
v1, v2, v3
sherry
v1, v2
yujunz
v1
The byx/v1 split contains 140 unique prompts and 300 generated images. For… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.enni-child-speech-synthesislicense: mit
task_categories:
text-to-speech
automatic-speech-recognition
language:
en
tags:
speech
audio
child-speech
talkbank
size_categories:
10K<n<100K
TalkBank Child Speech Synthesis Dataset (Seed 1)
This dataset contains child speech synthesis data generated from the TalkBank FASA ENNI corpus.
Dataset Information
Number of Samples: 10032
Seed: 1
Audio Format: WAV (16kHz)
Source: TalkBank FASA ENNI
Data Structure
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/jsun39/enni-child-speech-synthesis.sd_asr_synthesis_datastortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.goormaghtigh-equation-frontier-synthesis-v11
Goormaghtigh Equation Frontier Synthesis v11.0.0
Two-Chart Reconstruction, Radical Bounds, Exact Fibres, and Proof-Carrying Computation
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease date: 28 September 2026Status: Partial result. The unrestricted Goormaghtigh conjecture is not proved.Verification status: all declared computational scopes in this release are replayable from source; the work is not externally peer reviewed or proof-assistant… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/goormaghtigh-equation-frontier-synthesis-v11.common_voice_17_0_romanian_speech_synthesisromanian_speech_synthesis_0_8_1\
The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.sd_asr_synthesis_data_v0_less_silenceRecursive-Task-Synthesis-Copy
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Recursive-Task-Synthesis-Copy.Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modelfacial-defect-synthesis
Facial Defect Synthesis — output/
PRIVATE research dataset (internal use / cross-machine transfer). Two parts:
synthetic/ — images generated with OpenAI gpt-image-2. Depict no real patients;
safe to share. Distribution in synthetic/DISTRIBUTION.md.
real/<disease>/ — third-party reference images collected from public sources to
guide synthesis. Per-folder _sources.csv records each image's source URL + license. Many
are restrictive (e.g. CC BY-NC-ND) and depict real, identifiable… See the full description on the dataset page: https://huggingface.co/datasets/WilliamQiu123/facial-defect-synthesis.common_voice_16_1_romanian_speech_synthesistest
Image Classification Statistics
Counts
Category
Chinese
English
Total
ChatGPT
100
98
198
Gemini
8
15
23
Seed
23
31
54
Total
131
144
275
Notes
ChatGPT/English currently contains 98 images. Based on the expected numbering from 001e to 100e, 2 images are missing.
The missing files are ChatGPT Image009e.png and ChatGPT Image011e.png.
According to the current project record, these two missing images were not produced because of… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/test.codenib-synthesis
CodeMiner Synthesis
Status: In active development. This dataset is an early work-in-progress.
Both the set of instances and the per-instance query catalog are growing, and
the schema may evolve. Counts shown below describe the current snapshot
only — they are not a final target.
A growing collection of LLM-synthesized natural-language code-search
evaluation queries, each grounded on a real code symbol from a SWE-bench instance.
Design discussion and progress tracking:… See the full description on the dataset page: https://huggingface.co/datasets/sysevol-ai/codenib-synthesis.Neural-Solver-Synthesis-Final-Evidence-v1
Neural Solver Synthesis Final Evidence v1
This repository is the immutable large-artifact companion to the public code
release for "Beyond Inference-Time Search: Reinforcement Learning Synthesizes
Reusable Solvers."
Public code anchor:
07a798e7d7eca736cd1ef13a15209d402d401ef6
Current public release:
Neural Solver Synthesis evidence snapshot v1.0.2
Interactive experiment companion:
Neural Solver Synthesis
Aggregate evidence snapshot:
Final Evidence
Public policy checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1.Force-Controlled-Robotic-Mechanochemical-Synthesiscommon_voice_romanian_speech_synthesisChinese_Male_Speech_Synthesis_Corpus_Live_Streaming_for_Sales
ID
King-TTS-272
Duration
4.32 hours
Language
Chinese
URL
https://dataoceanai.com/datasets/tts/chinese-male-speech-synthesis-corpus-live-streaming-for-sales/
Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles
ID
King-TTS-241
Duration
8.56 hours
Speakers
100 People
Labeling Details
Pronunciation, Rhythm, Breath sounds marked with {hx}
Language
Chinese
Description
Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods.
URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles.
