datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.structured_data_merged_v2v5_0222
Dataset Card for structured_data_merged_v2v5_0222
Dataset Details
Dataset Description
structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV.
It was created by deduplicating and merging the following two existing datasets:
u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.combined-structured-dataset
Combined Structured Output Dataset
このデータセットは、構造化出力生成タスクのための統合データセットです。
Dataset Details
Total Samples: 3,425
Source Datasets:
u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples)
daichira/structured-3k-mix-sft (3,000 samples)
Preprocessing:
Format normalization to 3-turn (system/user/assistant)
Deduplication by user content
Quality filtering
Format Distribution
JSON: 685 (20.0%)
YAML: 485 (14.2%)
TOML: 685 (20.0%)
XML: 885 (25.8%)
CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.
