datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.synthesis
NuBerea/synthesis
A cross-corpus synthesis layer for the study of early Jewish and Christian literature.
Each config joins pericope-level text units from one corpus — the canonical Bible
(Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag
Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted
from those units and with links into a shared concept vocabulary. The result is a set
of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.Heterogenous_Synthesis_Benchmark
Heterogenous_Synthesis_Benchmark
This repository presents a diverse tabular data generation benchmark. We invite you to refer to our paper on arxiv to explore the mechanism behind our data diversity, which we called Distribution-Guided-Rule (DGR). Within this benchmark, you can experience how diverse preference data coverage combined with customized generation enhances post-training performance.
Additionally, Heterogenous_Synthesis_Benchmark includes a comprehensive toolkit for… See the full description on the dataset page: https://huggingface.co/datasets/CurryOvO/Heterogenous_Synthesis_Benchmark.literary-synthesis
Literary Synthesis
This dataset repurposes the original agentlans/literary-reasoning
data by reformatting it as creative writing prompts paired with literary-style outputs.
Writing style attributes were put in random order, with prompts randomly either prepended or appended.
The output text has been cleaned to make it suitable for creative writing and literary generation tasks.
The rows were sorted by increasing reading difficulty for curriculum learning.
organic-chemistry-synthesis-planning-corpus
Organic Chemistry Synthesis Planning Corpus
Status: actively ingesting. A comprehensive reaction backbone is already uploaded
(millions of reactions; see Ingested data below). Curation and
additional sources are ongoing. See Roadmap.
Quickstart (for students / first-time users)
You need a free Hugging Face account, and to accept this dataset's terms on its page (it's gated).
pip install datasets transformers
huggingface-cli login
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/organic-chemistry-synthesis-planning-corpus.Synthesis-Properties-Database-for-NanomaterialsThis is a database for nanomaterial synthesis. By fine-tuning the qwen3-14b model, it can extract synthesis steps, synthesis routes, and corresponding product properties (primarily including size, morphology, absorption spectra, and emission spectra) from target paragraphs. This fine-tuning project can be found at https://github.com/ime1452/Synthesis-Properties-Database-for-Nanomaterials.
The database contains two files: dataset.json holds the raw, unprocessed data, while… See the full description on the dataset page: https://huggingface.co/datasets/Kai-gu/Synthesis-Properties-Database-for-Nanomaterials.movement-synthesis-dataset
Movement Data Synthesis Dataset
Dataset Summary
This dataset contains 106 examples of movement tracking data specifically designed for training Large Language Models to generate synthetic physiotherapy and rehabilitation movement data. The dataset focuses on left arm circular exercises performed in a clockwise direction, captured using MediaPipe pose estimation technology.
Intended Use
Primary Use Cases
Fine-tuning LLMs for synthetic movement data… See the full description on the dataset page: https://huggingface.co/datasets/lucasbrandao/movement-synthesis-dataset.total_synthesis_of_new_drugs
Total Synthesis of New Drugs
This dataset contains basic information and raw synthesis route data for 60 drugs, along with 600 multimodal reasoning supervision (CoT) fine-tuning data.
The experimental and computational work in this dataset run on the Huawei Cloud AI Compute Service. We appreciate the stable compute supply from this platform.
新药化学全合成路线
该数据集包含 60 种药物的基本信息与合成路线原始数据,以及 600 条多模态思维监督微调数据。
本数据集的实验与计算工作依托于华为昇腾AI云服务平台完成,特此对其提供的稳定算力支持表示感谢。
Data-Synthesis-422K
Veri Setleri Hakkında / About the Datasets
Bu dosya, çeşitli veri setlerinin özelliklerini ve kullanım alanlarını özetlemektedir. / This document summarizes the features and use cases of various datasets.
anthracite-org/kalo-opus-instruct-22k-no-refusal
Açıklama / Description: Bu veri seti, çeşitli talimat ve yanıt çiftlerini içeren geniş bir koleksiyondur. Eğitim ve değerlendirme süreçlerinde kullanılmak üzere tasarlanmıştır. / This dataset contains a large collection… See the full description on the dataset page: https://huggingface.co/datasets/Kasimyildirim/Data-Synthesis-422K.resiplus-synthesis-dataset
ResiPlus Synthesis Dataset
Response synthesis dataset for nursing home assistant. Generates professional, structured responses from SQL and vector search results.
Dataset Details
Examples: 600
Language: Spanish (es)
Format: Chat messages (system, user, assistant)
Use Case: Fine-tuning LLMs for nursing home management system
Usage
from datasets import load_dataset
dataset = load_dataset("Alejandro284/resiplus-synthesis-dataset")
Training with HF… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro284/resiplus-synthesis-dataset.Final_Answer_Synthesis_with_Provenance
🇰🇿 Kazakh Final Answer Synthesis and Multi-Tool Task Automation Dataset
Dataset Summary
Kazakh Final Answer Synthesis and Multi-Tool Task Automation Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require multi-tool execution and final answer synthesis.
The dataset contains user requests, available tool schemas, expected tool calls, simulated tool outputs, and full multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Final_Answer_Synthesis_with_Provenance.
