datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen-Terminal-ToolBench-Processed-Tokenized
Qwen Terminal ToolBench Processed Datasets
Qwen-family processed/template-applied and selected tokenized terminal datasets.
Contents
qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text
qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text
qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels
qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.sft_processed_large_split
sft_processed_large — profile-disjoint split
This is the train / val / test split of Xuhui/sft_processed_large, the
OdysSim midtraining corpus (21.4M interactions across 63 datasets).
Split structure
split
rows
how it's built
train
21.20M
what's left after val + test are carved out
val
28K
per-dataset random sample, in-distribution; for checkpoint selection
test
128K
profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.ShareGPT-Processed
ShareGPT-Processed
The RyokoAI/ShareGPT52K dataset, converted to Markdown and labeled with the language used.
Acknowledgements
vinta/pangu.js — To insert whitespace between CJK (Chinese, Japanese, Korean) and half-width characters (alphabetical letters, numerical digits and symbols).
matthewwithanm/python-markdownify — Provides a starting point to convert HTML to Markdown.
BYVoid/OpenCC — Conversions between Traditional Chinese and Simplified Chinese.
aboSamoor/polyglot… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/ShareGPT-Processed.Llama-HybridDiffusion-processed-data-run1
Llama-HybridDiffusion processed training mixture — run 1
Built with Llama.
This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot
used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names,
dataset_info.json, state.json, and Arrow shard boundaries are retained so the data
can be downloaded and supplied to the existing training configuration without a lossy
format conversion.
Exact snapshot inventory
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.code_contest_processed
Dataset Card for Code Contest Processed
Dataset Summary
This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source.
Columns Description
id : unique string associated with a problem
description : problem description
code : one correct code for the problem
language : programming language used for code
test_samples : contains inputs and their… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_processed.SaaS-ProcessTwin
SaaS-ProcessTwin
Connected multilingual SaaS process simulations for causal decision reasoning.
SaaS-ProcessTwin is a synthetic benchmark of connected SaaS customer-risk cases. Each case is generated around a hidden object-centric event ledger and then projected into multilingual customer tickets, support notes, CRM summaries, incident updates, belief states, decisions, consequences, and counterfactual branches.
Models are evaluated on process reconstruction, belief tracking… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/SaaS-ProcessTwin.yeji-processed
██████╗ ██████╗ ██████╗ ██████╗███████╗███████╗███████╗███████╗██████╗
██╔══██╗██╔══██╗██╔═══██╗██╔════╝██╔════╝██╔════╝██╔════╝██╔════╝██╔══██╗
██████╔╝██████╔╝██║ ██║██║ █████╗ ███████╗███████╗█████╗ ██║ ██║
██╔═══╝ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║╚════██║██╔══╝ ██║ ██║
██║ ██║ ██║╚██████╔╝╚██████╗███████╗███████║███████║███████╗██████╔╝
╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚═════╝╚══════╝╚══════╝╚══════╝╚══════╝╚═════╝
⚡ REFINED TRAINING DATA ⚡
>… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-processed.OMB-Circular-A11-Section-120-Apportionment-Process
Dataset Description
The OMB Circular A-11 Section 120 Apportionment Process Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the federal apportionment process administered by the Office of Management and Budget.
The dataset was developed from Section 120, “Apportionment Process,” of OMB Circular No. A-11, Preparation, Submission, and Execution of the Budget. Section 120 is part of the Circular’s budget-execution… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A11-Section-120-Apportionment-Process.The_Stack_Processed-v2
🔥 The Stack Processed V2
A curated, balanced, and ML-optimized multi-language programming dataset
🎯 Why Choose This Dataset?
A meticulously curated version of "The Stack" optimized for training robust multi-language code models. Perfect balance between quality, diversity, and usability.
✨ Key Advantages:
🎯 Perfect Balance: ~10,000 files per major programming language
⚡ Training-Ready: Parquet format optimized for ML workflows
🏆 Superior Quality: 91.3% syntax… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/The_Stack_Processed-v2.nepali-recipes-qwen-processed
Nepali Recipes for Qwen Fine-tuning
Dataset Description
This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format.
Train Split: 900 recipes
Test Split: 327 recipes
Language: Nepali (ne)
Format: Qwen ChatML
Base Model: Qwen/Qwen2-1.5B
Dataset Structure
Data Fields
text: Full ChatML formatted prompt with answer (for training)
test_text: ChatML prompt without answer (for inference)
name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.tw-processed-law-article
Dataset Card for tw-processed-law-article
tw-processed-law-article 是一個中華民國法規條文之結構化資料集,以條文為單位展開,包含 230,974 筆條文,涵蓋 11,462 部不同法規,橫跨憲法、法律與命令三個層級。每筆資料包含法規名稱、層級、條文內容、廢止註記與最後修正日期等欄位,適用於法律檢索系統、條文問答模型,或作為其他法律衍生資料集之結構化底層語料。
Dataset Details
Dataset Description
本資料集整理自中華民國全國法規資料庫(law.moj.gov.tw)之公開法規條文。原始法規資料經處理後以單條條文為一筆資料(one row per article),每筆附帶法規名稱、層級分類、條文全文、廢止註記、最後修正日期與 API 更新日期等元資料。
層級分佈:
層級
筆數
說明
命令
180,757
由主管機關訂定之命令、規則、辦法等
法律
49,977
經立法院三讀通過之法律
憲法… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-article.processflow
ProcessFlow
A multi-format, process-centric code dataset for training LLM agents.
✅ EMPIRICALLY VALIDATED (2026-04-10). Fine-tuning Qwen2.5-1.5B base on
v1.7 (108K training samples, 3 epochs, LoRA r=32) produced a
+0.681 ProcessFlow-Eval delta (0.217 → 0.899) with no HumanEval
regression and PPL improvement of -4.62 nats on held-out test data.
All 3 validation gates passed decisively. See Empirical validation
section below. Trained adapter:… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/processflow.gold-silver-mineral-process-sft-candidates
Gold/silver mineral-process SFT candidates
English chat pairs about gold/silver and transferable hard-rock mineral processing. Each row uses a messages list (user, then assistant).
Filtered subset of public Hugging Face datasets. Not the original uploads.
Upstream
Rows here
Lyntas/mininggpt_training_dataset
14,752
polyhedralai/mining_concepts
125
Configs
Config
Rows
default
14,877
mininggpt_strict
14,752
mining_concepts_strict
125… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-sft-candidates.cleo-process-analytics-v1
Cleo Process Analytics v1
cleo-process-analytics-v1 is a 260-example SQL analytics dataset built for process-heavy analyst
workflows. The questions are designed to require multi-step SQL behavior such as joins, aggregations,
rankings, CTEs, windows, and occasional semantic-view use, while keeping answers deterministic and
execution-verified.
This dataset was created for the Cleo SQL analyst project:
github.com/Dreeseaw/cleo.
Contents
path
description… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-process-analytics-v1.medical-llm-finetuning-alignment-processed-datasetclinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.WebInstruct-Verified-Processed
WebInstruct-Verified-Processed
The WebInstruct-Verified-Processed dataset is used in the paper Characterizing, Evaluating, and Optimizing Complex Reasoning. It is a processed version of WebInstruct-verified, formatted for RL training with verifiable rewards. In the paper, this dataset is used as the RL prompt/data source for TRM-guided reinforcement learning optimization, where models are trained with rule-based verifiers and an auxiliary Thinking Reward Model (TRM) signal.… See the full description on the dataset page: https://huggingface.co/datasets/zzzhr97/WebInstruct-Verified-Processed.gold-silver-mineral-process-cpt-candidates
Gold/silver mineral-process CPT candidates
English raw documents (text) about gold/silver and transferable hard-rock mineral processing.
Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only.
These are documents, not chat pairs.
Configs
Config
Rows
Notes
default
59,749
all English bands
english_high
18,864
publisher quality 4.00–4.59
english_middle
34,601
publisher quality 3.00–4.00
english_low
6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.process_comp_pool
Dataset Card: Synthetic Emotion Component Process Item Pool
1. Dataset Description
1.1. Dataset Summary
This dataset provides a large-scale, synthetically generated pool of self-report questionnaire items designed to measure different components of emotion. Traditional psychometric scale development is often bottlenecked by the initial item generation phase, which relies heavily on subjective subject-matter expert (SME) brainstorming. This dataset explores the… See the full description on the dataset page: https://huggingface.co/datasets/christiqn/process_comp_pool.claude-sonnet-4.6-processed-reasoningIn this dataset, Claude reasoning summaries were processed by Gemma-4-31B into more genuine reasoning like you might find on open-weight reasoning models. This won't recover the original thinking, but this should help improve training convergence when paired with other reasoning datasets.
Thanks to NVIDIA NIM and Parasail. And especially thanks big thanks to https://huggingface.co/datasets/Roman1111111/claude-sonnet-4.6-120000x
Note that the vast majority of the reasoning was constructed with… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/claude-sonnet-4.6-processed-reasoning.tw-processed-judgments-14B
Dataset Card for tw-processed-judgments-14B
tw-processed-judgments-14B 是一個中華民國司法判決書之大規模預訓練語料,收錄自 1996 年 1 月至最新之各審級判決書,總計約 14B tokens(約 40 GB,估計約 2,300 萬筆判決)。資料經專業法律領域知識與自然語言處理方法共同清理,移除裁定、簡易庭及格式不符之判決,僅保留具完整主文、事實、理由結構之判決書,適用於繁體中文法律領域 LLM 之預訓練與持續預訓練(CPT)。
Dataset Details
Dataset Description
原始司法院判決書公開文本存在大量格式雜訊、錯誤內容與缺漏段落,導致許多以此為語料訓練之台灣本地 LLM 雖接觸過大量判決書,仍無法有效提升法律領域之推理能力。本資料集針對此問題進行大量清理與篩選:
移除裁定類型(僅保留「判決」);
移除簡易庭判決;
篩選符合標準格式之判決書(包含主文、事實、理由三大段落);… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-judgments-14B.qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed
Qwen3 4B Thinking SFT v54 Processed Training View
This dataset is the processed and filtered training view used by the v54
Qwen3-4B-Thinking SFT recipe. It starts from
eewer/swerebench-traces-raw-source-targeted-limitations-compaction-full-20260616-2030 and uses the strict-passed raw2030 mini-swe
aligned view.
Rows are compressed JSONL.zst files under data/. Each row contains a
top-level messages column, optional tools, and scalar source mapping fields
such as source_uuid… See the full description on the dataset page: https://huggingface.co/datasets/eewer/qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed.tw-processed-law-ctx
Dataset Card for tw-processed-law-ctx
tw-processed-law-ctx 是一個中華民國(台灣)法規全文之後處理資料集,合計 11,462 筆。相較於 tw-processed-law-article(以單一條文為單位),本資料集將「同一法規之所有條文」合併為單一 text,提供整部法規之完整上下文,適合作為繁中法律 LLM 之持續預訓練(CPT)素材。
Dataset Details
Dataset Description
以「條文為單位」之語料雖便於精確查詢,但模型在訓練時難以學到整部法規之章節結構與條文之間之關聯。本資料集將每一部法規之所有條文依序合併為單一長文本,保留:
法規名稱(含章節標題);
所有條文之條號與本文;
該法規最近修正日期與 API 更新日期;
若法規已廢止,額外提供 abandon_note 欄位。
每筆亦同時附帶 level(法規位階:憲法 / 法律 / 命令 / 廢止)作為篩選依據。
Curated by: Liang Hsun… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-ctx.rpguild_processedpreprocessed version of chargoddard/rpguild into a fun little prompt format for finetuning
en-edgar-processed
SEC EDGAR filings (processed)
🌐 The Fin AI
Pretraining / reference corpus released by The Fin AI. Source: U.S. SEC EDGAR filings — https://www.sec.gov/search-filings.
Note: Some parquet files in this repository are empty (0 bytes).
Source
Text extracted from public filings on the SEC EDGAR system.
Structure
Rows: 2,839,488
Columns: text
Quick Start
from datasets import load_dataset
ds = load_dataset("TheFinAI/en-edgar-processed"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/en-edgar-processed.gr_processed_datasets
Greek processed corpora
🌐 The Fin AI
Processed Greek corpora (academic articles, company filings, general text, laws and regulations) as JSONL.
Task
pretraining corpus
Language
el
License
cc-by-4.0
Hindi-Marathi-Synonyms
Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह)
Overview
This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications.
The dataset provides word-synonym pairs that can be used for tasks like:
Semantic analysis
Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.roleplay_boundaries_processedProcessed version of limloop/roleplay_knowledge_boundaries.
Each sample is a chat in the messages format: a system persona description
(role, knowledge domains, taboo topics, communication style, evasion techniques)
followed by alternating user / assistant turns. The system message also asks
for short, in-character replies.
messages: list of {"role", "content"}
language: en or ru
Train split: 3771 rows.
docs-asof-processed
profgabrielramos/docs-asof-processed
Objetivo
Dataset textual preparado para auditoria de qualidade, limpeza reprodutível e uso em pipelines de RAG.
Origem
Documentos publicados em profgabrielramos/docs-asof-processed.
Versão
v1.0.1
Estatísticas públicas (auditoria)
Documentos: 803
Coluna textual: text
Duplicação normalizada: 2.24%
Nulos na coluna textual: 0.00%
Tamanho médio (chars): 1579.1
p50/p90 (chars): 502.0 / 2221.6
Idioma… See the full description on the dataset page: https://huggingface.co/datasets/profgabrielramos/docs-asof-processed.Language_Identification_v1
Dataset Card for Language Identification Dataset
Dataset Summary
A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications.
Languages and Distribution
Language Distribution:
Urdu 1000
Hindi 1000
Odia 1000
Tamil 1000
Kannada 1000
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.
