Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TMaxxx /agent-task-recursive-task-synthesis Apptainer pool for hamishivi/agent-task-recursive-task-synthesis This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. Apptainer images The pool currently contains 29,501 / 29… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-recursive-task-synthesis.5 likes16k downloads22d agoHugging Face02hamishivi /agent-task-recursive-task-synthesis Recursive-Task-Synthesis for tmax Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments. All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format. The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.text10K<n<100K0 likes2.7k downloads29d agoHugging Face03Zhongzhi1228 /Recursive-Task-Synthesis Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K19 likes2.4k downloads2mo agoHugging Face04Xnhyacinth /LongWorld-Synthesis-Workspace LongWorld synthesis workspace Machine-migration bundle for continuing synthesis. Train-ready SFT rows are in Xnhyacinth/LongWorld-Real-Workflows. Probe signing keys and credentials are excluded from this public dataset. Verification keys must be provisioned separately through the owning workspace. Layout Path What it is source_inventory/ HMAC-signed source records plus retained source bytes (arXiv tarballs, filings). Cannot be rebuilt without the same… See the full description on the dataset page: https://huggingface.co/datasets/Xnhyacinth/LongWorld-Synthesis-Workspace.0 likes1.3k downloads2d agoHugging Face05PrimeIntellect /Recursive-Task-Synthesis Recursive Task Synthesis Tasks without completed platform artifacts or with unresolved VM validation failures are temporarily excluded. exclusions.json records the exact IDs, reasons, build IDs where available, and evidence dates/runs. Exclusions affect both metadata rows and complete TAR task packages. Runtime failures are not image-build failures or proof of incorrect gold solutions. This filter does not establish that every retained task passes gold validation. Restore a task… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K3 likes1.1k downloads12d agoHugging Face06open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes833 downloads21d agoHugging Face07Zhongzhi1228 /Recursive-Task-Synthesis-Trajectories Recursive Task Synthesis Trajectories This dataset contains 327,189 completed agent trajectories collected on recursively synthesized command-line tasks. Public identifiers are opaque and stable. The trajectory JSON retains messages, actions, observations, and token counts. Token-level log-probability arrays and duplicated debug/session captures are excluded from the public packages. metadata/trajectories.parquet: searchable trajectory metadata. metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.tabularreinforcement-learning100K<n<1M4 likes808 downloads2mo agoHugging Face08Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes561 downloads1mo agoHugging Face09Aviv-anthonnyolime /SIWIS_French_Speech_Synthesis_Database SIWIS French Speech Synthesis Database This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section. The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose. For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.audioautomatic-speech-recognition10K<n<100K0 likes530 downloads2y agoHugging Face10NuBerea /synthesisgated NuBerea/synthesis A cross-corpus synthesis layer for the study of early Jewish and Christian literature. Each config joins pericope-level text units from one corpus — the canonical Bible (Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted from those units and with links into a shared concept vocabulary. The result is a set of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.tabulartext-generation10K<n<100K0 likes484 downloads3d agoHugging Face11knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2, v3 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images.… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes470 downloads4d agoHugging Face12Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes439 downloads6mo agoHugging Face13Zhongzhi1228 /Recursive-Task-Synthesis-Quality-1K Recursive Task Synthesis Quality 1K This dataset contains 1,000 quality-selected, validated command-line task instances. It is a curated subset of the Recursive Task Synthesis dataset. Public task and group identifiers are opaque and stable across both datasets. Selection The subset was selected from 37,484 validated tasks using structural and safety checks, two-pass semantic review, strict gates for instruction clarity, instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.tabularreinforcement-learning1K<n<10K0 likes349 downloads2mo agoHugging Face14jsun39 /enni-child-speech-synthesislicense: mit task_categories: text-to-speech automatic-speech-recognition language: en tags: speech audio child-speech talkbank size_categories: 10K<n<100K TalkBank Child Speech Synthesis Dataset (Seed 1) This dataset contains child speech synthesis data generated from the TalkBank FASA ENNI corpus. Dataset Information Number of Samples: 10032 Seed: 1 Audio Format: WAV (16kHz) Source: TalkBank FASA ENNI Data Structure The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/jsun39/enni-child-speech-synthesis.audio1K<n<10K13 likes238 downloads6mo agoHugging Face15frankie137 /sd_asr_synthesis_datatabularn<1K0 likes204 downloads4mo agoHugging Face16PureOne /goormaghtigh-equation-frontier-synthesis-v11 Goormaghtigh Equation Frontier Synthesis v11.0.0 Two-Chart Reconstruction, Radical Bounds, Exact Fibres, and Proof-Carrying Computation Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease date: 28 September 2026Status: Partial result. The unrestricted Goormaghtigh conjecture is not proved.Verification status: all declared computational scopes in this release are replayable from source; the work is not externally peer reviewed or proof-assistant… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/goormaghtigh-equation-frontier-synthesis-v11.0 likes195 downloads12d agoHugging Face17xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K10 likes170 downloads3y agoHugging Face18gigant /romanian_speech_synthesis_0_8_1\ The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.automatic-speech-recognition14 likes161 downloads4y agoHugging Face19VladS159 /common_voice_17_0_romanian_speech_synthesisaudio10K<n<100K8 likes161 downloads2y agoHugging Face20WilliamQiu123 /facial-defect-synthesis Facial Defect Synthesis — output/ PRIVATE research dataset (internal use / cross-machine transfer). Two parts: synthetic/ — images generated with OpenAI gpt-image-2. Depict no real patients; safe to share. Distribution in synthetic/DISTRIBUTION.md. real/<disease>/ — third-party reference images collected from public sources to guide synthesis. Per-folder _sources.csv records each image's source URL + license. Many are restrictive (e.g. CC BY-NC-ND) and depict real, identifiable… See the full description on the dataset page: https://huggingface.co/datasets/WilliamQiu123/facial-defect-synthesis.imagen<1K0 likes140 downloads3mo agoHugging Face21VladS159 /common_voice_16_1_romanian_speech_synthesisaudio10K<n<100K9 likes133 downloads3y agoHugging Face22frankie137 /sd_asr_synthesis_data_v0_less_silencetabularn<1K0 likes124 downloads4mo agoHugging Face23VladS159 /common_voice_romanian_speech_synthesisaudio10K<n<100K7 likes120 downloads3y agoHugging Face24reshinthadith /synthetic_program_synthesis_python_1Mtext100K<n<1M9 likes114 downloads4y agoHugging Face25DataoceanAI /Chinese_Male_Speech_Synthesis_Corpus_Live_Streaming_for_Sales ID King-TTS-272 Duration 4.32 hours Language Chinese URL https://dataoceanai.com/datasets/tts/chinese-male-speech-synthesis-corpus-live-streaming-for-sales/ 7 likes110 downloads2y agoHugging Face26DTLR /concepticon-wiktionary-synthesis Concepticon-Wiktionary Synthesis: Dataset for Tagging the EFEO-CNRS-SOAS Lexicon CHAPTER I. INTRODUCTION CHAPTER II. FIELDWORK AND COLLECTION STRATEGY CHAPTER III. TECHNICAL METHODOLOGY &nbsp;&nbsp;A. Dataset Synthesis &nbsp;&nbsp;B. Dataset Structure REFERENCES Abstract Abstract A synthesised, multilingual conceptual alignment dataset is presented to facilitate a shared task aimed at the semantic tagging of the EFEO-CNRS-SOAS lexicon with abstract… See the full description on the dataset page: https://huggingface.co/datasets/DTLR/concepticon-wiktionary-synthesis.tabulartoken-classification1K<n<10K1 likes106 downloads18h agoHugging Face27DataoceanAI /Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles ID King-TTS-241 Duration 8.56 hours Speakers 100 People Labeling Details Pronunciation, Rhythm, Breath sounds marked with {hx} Language Chinese Description Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods. URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles.10 likes104 downloads2y agoHugging Face28DataoceanAI /Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales ID King-TTS-271 Duration 4.24 hours Language Chinese URL https://dataoceanai.com/datasets/tts/chinese-female-speech-synthesis-corpus-live-streaming-for-sales/ 9 likes103 downloads2y agoHugging Face29FForty7 /Force-Controlled-Robotic-Mechanochemical-Synthesis3dn<1K0 likes102 downloads4mo agoHugging Face30DataoceanAI /American_English_Male_Speech_Synthesis_Corpus_Gentle_and_Mature_Aged_30_40 ID King-TTS-286 Duration 3.02 hours Language English Labeled Details Multi-emotion - Neutral, Happy, Angry, Sad, Shocked, Hateful, Scared, Shouting, Crying, Laughing, Weak URL https://dataoceanai.com/datasets/asr/indonesian-speech-recognition-corpus/ 10 likes99 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.