Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hamishivi /agent-task-recursive-task-synthesis Recursive-Task-Synthesis for tmax Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments. All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format. The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.text10K<n<100K0 likes2.7k downloads29d agoHugging Face02Zhongzhi1228 /Recursive-Task-Synthesis Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K19 likes2.4k downloads2mo agoHugging Face03PrimeIntellect /Recursive-Task-Synthesis Recursive Task Synthesis Tasks without completed platform artifacts or with unresolved VM validation failures are temporarily excluded. exclusions.json records the exact IDs, reasons, build IDs where available, and evidence dates/runs. Exclusions affect both metadata rows and complete TAR task packages. Runtime failures are not image-build failures or proof of incorrect gold solutions. This filter does not establish that every retained task passes gold validation. Restore a task… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K3 likes1.1k downloads12d agoHugging Face04open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes833 downloads21d agoHugging Face05Zhongzhi1228 /Recursive-Task-Synthesis-Trajectories Recursive Task Synthesis Trajectories This dataset contains 327,189 completed agent trajectories collected on recursively synthesized command-line tasks. Public identifiers are opaque and stable. The trajectory JSON retains messages, actions, observations, and token counts. Token-level log-probability arrays and duplicated debug/session captures are excluded from the public packages. metadata/trajectories.parquet: searchable trajectory metadata. metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.tabularreinforcement-learning100K<n<1M4 likes808 downloads2mo agoHugging Face06Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes561 downloads1mo agoHugging Face07Aviv-anthonnyolime /SIWIS_French_Speech_Synthesis_Database SIWIS French Speech Synthesis Database This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section. The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose. For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.audioautomatic-speech-recognition10K<n<100K0 likes530 downloads2y agoHugging Face08NuBerea /synthesisgated NuBerea/synthesis A cross-corpus synthesis layer for the study of early Jewish and Christian literature. Each config joins pericope-level text units from one corpus — the canonical Bible (Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted from those units and with links into a shared concept vocabulary. The result is a set of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.tabulartext-generation10K<n<100K0 likes484 downloads3d agoHugging Face09knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2, v3 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images.… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes470 downloads4d agoHugging Face10Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes439 downloads6mo agoHugging Face11Zhongzhi1228 /Recursive-Task-Synthesis-Quality-1K Recursive Task Synthesis Quality 1K This dataset contains 1,000 quality-selected, validated command-line task instances. It is a curated subset of the Recursive Task Synthesis dataset. Public task and group identifiers are opaque and stable across both datasets. Selection The subset was selected from 37,484 validated tasks using structural and safety checks, two-pass semantic review, strict gates for instruction clarity, instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.tabularreinforcement-learning1K<n<10K0 likes349 downloads2mo agoHugging Face12frankie137 /sd_asr_synthesis_datatabularn<1K0 likes204 downloads4mo agoHugging Face13xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K10 likes170 downloads3y agoHugging Face14VladS159 /common_voice_17_0_romanian_speech_synthesisaudio10K<n<100K8 likes161 downloads2y agoHugging Face15VladS159 /common_voice_16_1_romanian_speech_synthesisaudio10K<n<100K9 likes133 downloads3y agoHugging Face16frankie137 /sd_asr_synthesis_data_v0_less_silencetabularn<1K0 likes124 downloads4mo agoHugging Face17VladS159 /common_voice_romanian_speech_synthesisaudio10K<n<100K7 likes120 downloads3y agoHugging Face18reshinthadith /synthetic_program_synthesis_python_1Mtext100K<n<1M9 likes114 downloads4y agoHugging Face19DTLR /concepticon-wiktionary-synthesis Concepticon-Wiktionary Synthesis: Dataset for Tagging the EFEO-CNRS-SOAS Lexicon CHAPTER I. INTRODUCTION CHAPTER II. FIELDWORK AND COLLECTION STRATEGY CHAPTER III. TECHNICAL METHODOLOGY &nbsp;&nbsp;A. Dataset Synthesis &nbsp;&nbsp;B. Dataset Structure REFERENCES Abstract Abstract A synthesised, multilingual conceptual alignment dataset is presented to facilitate a shared task aimed at the semantic tagging of the EFEO-CNRS-SOAS lexicon with abstract… See the full description on the dataset page: https://huggingface.co/datasets/DTLR/concepticon-wiktionary-synthesis.tabulartoken-classification1K<n<10K1 likes106 downloads18h agoHugging Face20FForty7 /Force-Controlled-Robotic-Mechanochemical-Synthesis3dn<1K0 likes102 downloads4mo agoHugging Face21JojoZhu /cosmos-trajectory-synthesisgated COSMOS Synthetic Traffic Trajectory Dataset Related releases: Earlier 1042-scene export (incl. normal split) · Real COSMOS trajectories Description Synthetic multi-agent traffic trajectories at a fixed urban intersection, generated by the COSMOS pipeline with Protocol V2 LLM backends and Tier-1 quality gates (geometry + kinematics). Each scene is created from a natural language prompt and processed through: Scene Planner → GMM Sampling → Waypoint Filter → Event… See the full description on the dataset page: https://huggingface.co/datasets/JojoZhu/cosmos-trajectory-synthesis.tabularothern<1K1 likes94 downloads3d agoHugging Face22sysevol-ai /codenib-synthesis CodeMiner Synthesis Status: In active development. This dataset is an early work-in-progress. Both the set of instances and the per-instance query catalog are growing, and the schema may evolve. Counts shown below describe the current snapshot only — they are not a final target. A growing collection of LLM-synthesized natural-language code-search evaluation queries, each grounded on a real code symbol from a SWE-bench instance. Design discussion and progress tracking:… See the full description on the dataset page: https://huggingface.co/datasets/sysevol-ai/codenib-synthesis.textn<1K3 likes90 downloads2mo agoHugging Face23CurryOvO /Heterogenous_Synthesis_Benchmark Heterogenous_Synthesis_Benchmark This repository presents a diverse tabular data generation benchmark. We invite you to refer to our paper on arxiv to explore the mechanism behind our data diversity, which we called Distribution-Guided-Rule (DGR). Within this benchmark, you can experience how diverse preference data coverage combined with customized generation enhances post-training performance. Additionally, Heterogenous_Synthesis_Benchmark includes a comprehensive toolkit for… See the full description on the dataset page: https://huggingface.co/datasets/CurryOvO/Heterogenous_Synthesis_Benchmark.texttext-classification10K<n<100K1 likes66 downloads6mo agoHugging Face24introvoyz041 /Force-Controlled-Robotic-Mechanochemical-Synthesis3dn<1K0 likes66 downloads5mo agoHugging Face25ngwgsang /vietquill-qcpg-100k-synthesis-questiontabular100K<n<1M1 likes65 downloads16d agoHugging Face26Nexdata-kr /2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus Description 48kHz, 24bit 품질의 영어 음성 데이터셋으로, 전문 녹음 스튜디오에서 전문 성우 2명(남성 1명, 여성 1명)의 음성을 수집했습니다. 주어진 주제에 대한 즉흥 발화, 다단계 감정, 단일 감정 및 준언어적 특성(Paralinguistic Features) 등 다양한 음성 콘텐츠를 포함합니다. 텍스트, 감정 및 준언어적 특성에 대한 어노테이션을 제공하며, 음성 합성(Speech Synthesis) 등의 음성 AI 모델 개발 및 학습에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/tts/1540?source=hf.kr Specifications Format 48kHz, 24bit, 비압축 WAV, 모노 채널 Recording Environment 전문 녹음 스튜디오… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus.audion<1K0 likes62 downloads16d agoHugging Face27AmanPriyanshu /Recursive-Task-Synthesis-Copy Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Recursive-Task-Synthesis-Copy.tabularreinforcement-learning10K<n<100K0 likes51 downloads2mo agoHugging Face28quocanh34 /soict_train_synthesis_dataset Dataset Card for "soict_train_synthesis_dataset" More Information needed text1K<n<10K0 likes43 downloads3y agoHugging Face29agentlans /literary-synthesis Literary Synthesis This dataset repurposes the original agentlans/literary-reasoning data by reformatting it as creative writing prompts paired with literary-style outputs. Writing style attributes were put in random order, with prompts randomly either prepended or appended. The output text has been cleaned to make it suitable for creative writing and literary generation tasks. The rows were sorted by increasing reading difficulty for curriculum learning. texttext-generation1K<n<10K3 likes42 downloads1y agoHugging Face30quocanh34 /synthesis_data_v1 Dataset Card for "synthesis_data_v1" More Information needed text1K<n<10K0 likes41 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.