Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLM-OS-Models /Qwen-Terminal-ToolBench-Processed-Tokenized Qwen Terminal ToolBench Processed Datasets Qwen-family processed/template-applied and selected tokenized terminal datasets. Contents qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.text-generation0 likes2.2k downloads4mo agoHugging Face02Xuhui /sft_processed_large_split sft_processed_large — profile-disjoint split This is the train / val / test split of Xuhui/sft_processed_large, the OdysSim midtraining corpus (21.4M interactions across 63 datasets). Split structure split rows how it's built train 21.20M what's left after val + test are carved out val 28K per-dataset random sample, in-distribution; for checkpoint selection test 128K profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.texttext-generation10M<n<100M1 likes1.4k downloads6mo agoHugging Face03zetavg /ShareGPT-Processed ShareGPT-Processed The RyokoAI/ShareGPT52K dataset, converted to Markdown and labeled with the language used. Acknowledgements vinta/pangu.js — To insert whitespace between CJK (Chinese, Japanese, Korean) and half-width characters (alphabetical letters, numerical digits and symbols). matthewwithanm/python-markdownify — Provides a starting point to convert HTML to Markdown. BYVoid/OpenCC — Conversions between Traditional Chinese and Simplified Chinese. aboSamoor/polyglot… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/ShareGPT-Processed.texttext-generation10K<n<100K30 likes289 downloads3y agoHugging Face04Arushhh /Llama-HybridDiffusion-processed-data-run1 Llama-HybridDiffusion processed training mixture — run 1 Built with Llama. This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names, dataset_info.json, state.json, and Arrow shard boundaries are retained so the data can be downloaded and supplied to the existing training configuration without a lossy format conversion. Exact snapshot inventory Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.text-generation1M<n<10M0 likes226 downloads2mo agoHugging Face05iamtarun /code_contest_processed Dataset Card for Code Contest Processed Dataset Summary This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem language : programming language used for code test_samples : contains inputs and their… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_processed.texttext-generation10K<n<100K3 likes165 downloads3y agoHugging Face06Lightcap /SaaS-ProcessTwin SaaS-ProcessTwin Connected multilingual SaaS process simulations for causal decision reasoning. SaaS-ProcessTwin is a synthetic benchmark of connected SaaS customer-risk cases. Each case is generated around a hidden object-centric event ledger and then projected into multilingual customer tickets, support notes, CRM summaries, incident updates, belief states, decisions, consequences, and counterfactual branches. Models are evaluated on process reconstruction, belief tracking… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/SaaS-ProcessTwin.tabularquestion-answering10M<n<100M1 likes112 downloads5mo agoHugging Face07tellang /yeji-processed ██████╗ ██████╗ ██████╗ ██████╗███████╗███████╗███████╗███████╗██████╗ ██╔══██╗██╔══██╗██╔═══██╗██╔════╝██╔════╝██╔════╝██╔════╝██╔════╝██╔══██╗ ██████╔╝██████╔╝██║ ██║██║ █████╗ ███████╗███████╗█████╗ ██║ ██║ ██╔═══╝ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║╚════██║██╔══╝ ██║ ██║ ██║ ██║ ██║╚██████╔╝╚██████╗███████╗███████║███████║███████╗██████╔╝ ╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚═════╝╚══════╝╚══════╝╚══════╝╚══════╝╚═════╝ ⚡ REFINED TRAINING DATA ⚡ >… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-processed.texttext-generation10K<n<100K0 likes91 downloads9mo agoHugging Face08leeroy-jankins /OMB-Circular-A11-Section-120-Apportionment-Process Dataset Description The OMB Circular A-11 Section 120 Apportionment Process Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the federal apportionment process administered by the Office of Management and Budget. The dataset was developed from Section 120, “Apportionment Process,” of OMB Circular No. A-11, Preparation, Submission, and Execution of the Budget. Section 120 is part of the Circular’s budget-execution… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A11-Section-120-Apportionment-Process.documentquestion-answeringn<1K1 likes87 downloads2mo agoHugging Face09vinsblack /The_Stack_Processed-v2 🔥 The Stack Processed V2 A curated, balanced, and ML-optimized multi-language programming dataset 🎯 Why Choose This Dataset? A meticulously curated version of "The Stack" optimized for training robust multi-language code models. Perfect balance between quality, diversity, and usability. ✨ Key Advantages: 🎯 Perfect Balance: ~10,000 files per major programming language ⚡ Training-Ready: Parquet format optimized for ML workflows 🏆 Superior Quality: 91.3% syntax… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/The_Stack_Processed-v2.tabulartext-generation100K<n<1M4 likes68 downloads1y agoHugging Face10sijanpaudel /nepali-recipes-qwen-processed Nepali Recipes for Qwen Fine-tuning Dataset Description This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format. Train Split: 900 recipes Test Split: 327 recipes Language: Nepali (ne) Format: Qwen ChatML Base Model: Qwen/Qwen2-1.5B Dataset Structure Data Fields text: Full ChatML formatted prompt with answer (for training) test_text: ChatML prompt without answer (for inference) name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.texttext-generation1K<n<10K0 likes67 downloads1y agoHugging Face11lianghsun /tw-processed-law-article Dataset Card for tw-processed-law-article tw-processed-law-article 是一個中華民國法規條文之結構化資料集,以條文為單位展開,包含 230,974 筆條文,涵蓋 11,462 部不同法規,橫跨憲法、法律與命令三個層級。每筆資料包含法規名稱、層級、條文內容、廢止註記與最後修正日期等欄位,適用於法律檢索系統、條文問答模型,或作為其他法律衍生資料集之結構化底層語料。 Dataset Details Dataset Description 本資料集整理自中華民國全國法規資料庫(law.moj.gov.tw)之公開法規條文。原始法規資料經處理後以單條條文為一筆資料(one row per article),每筆附帶法規名稱、層級分類、條文全文、廢止註記、最後修正日期與 API 更新日期等元資料。 層級分佈: 層級 筆數 說明 命令 180,757 由主管機關訂定之命令、規則、辦法等 法律 49,977 經立法院三讀通過之法律 憲法… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-article.texttext-generation100K<n<1M3 likes65 downloads6mo agoHugging Face12caiovicentino1 /processflow ProcessFlow A multi-format, process-centric code dataset for training LLM agents. ✅ EMPIRICALLY VALIDATED (2026-04-10). Fine-tuning Qwen2.5-1.5B base on v1.7 (108K training samples, 3 epochs, LoRA r=32) produced a +0.681 ProcessFlow-Eval delta (0.217 → 0.899) with no HumanEval regression and PPL improvement of -4.62 nats on held-out test data. All 3 validation gates passed decisively. See Empirical validation section below. Trained adapter:… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/processflow.text-generation100K<n<1M1 likes55 downloads6mo agoHugging Face13hicham-taoufik /gold-silver-mineral-process-sft-candidates Gold/silver mineral-process SFT candidates English chat pairs about gold/silver and transferable hard-rock mineral processing. Each row uses a messages list (user, then assistant). Filtered subset of public Hugging Face datasets. Not the original uploads. Upstream Rows here Lyntas/mininggpt_training_dataset 14,752 polyhedralai/mining_concepts 125 Configs Config Rows default 14,877 mininggpt_strict 14,752 mining_concepts_strict 125… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-sft-candidates.textquestion-answering10K<n<100K0 likes55 downloads10d agoHugging Face14dreeseaw /cleo-process-analytics-v1 Cleo Process Analytics v1 cleo-process-analytics-v1 is a 260-example SQL analytics dataset built for process-heavy analyst workflows. The questions are designed to require multi-step SQL behavior such as joins, aggregations, rankings, CTEs, windows, and occasional semantic-view use, while keeping answers deterministic and execution-verified. This dataset was created for the Cleo SQL analyst project: github.com/Dreeseaw/cleo. Contents path description… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-process-analytics-v1.text-generationn<1K0 likes48 downloads4mo agoHugging Face15dworsleytonks /medical-llm-finetuning-alignment-processed-datasettexttext-generation10K<n<100K0 likes43 downloads10mo agoHugging Face16ClarusC64 /clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1 Goal Test if a model can hold separate reasoning streams at once Detect constraint dismissal Detect bleed-over where one stream turns into claims in the other What it measures streams_heldResponse acknowledges and maintains both streams bleed_overConstraint stream improperly becomes a medical claim, or vice versa premature_synthesisResponse forces a single solution that silences one stream assumption_collapseResponse drops a premise entirely Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.texttext-generationn<1K0 likes41 downloads9mo agoHugging Face17zzzhr97 /WebInstruct-Verified-Processed WebInstruct-Verified-Processed The WebInstruct-Verified-Processed dataset is used in the paper Characterizing, Evaluating, and Optimizing Complex Reasoning. It is a processed version of WebInstruct-verified, formatted for RL training with verifiable rewards. In the paper, this dataset is used as the RL prompt/data source for TRM-guided reinforcement learning optimization, where models are trained with rule-based verifiers and an auxiliary Thinking Reward Model (TRM) signal.… See the full description on the dataset page: https://huggingface.co/datasets/zzzhr97/WebInstruct-Verified-Processed.texttext-generation100K<n<1M1 likes41 downloads4mo agoHugging Face18hicham-taoufik /gold-silver-mineral-process-cpt-candidates Gold/silver mineral-process CPT candidates English raw documents (text) about gold/silver and transferable hard-rock mineral processing. Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only. These are documents, not chat pairs. Configs Config Rows Notes default 59,749 all English bands english_high 18,864 publisher quality 4.00–4.59 english_middle 34,601 publisher quality 3.00–4.00 english_low 6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.tabulartext-generation100K<n<1M0 likes38 downloads16d agoHugging Face19christiqn /process_comp_pool Dataset Card: Synthetic Emotion Component Process Item Pool 1. Dataset Description 1.1. Dataset Summary This dataset provides a large-scale, synthetically generated pool of self-report questionnaire items designed to measure different components of emotion. Traditional psychometric scale development is often bottlenecked by the initial item generation phase, which relies heavily on subjective subject-matter expert (SME) brainstorming. This dataset explores the… See the full description on the dataset page: https://huggingface.co/datasets/christiqn/process_comp_pool.texttext-classification10K<n<100K1 likes31 downloads7mo agoHugging Face20MasonMac /claude-sonnet-4.6-processed-reasoningIn this dataset, Claude reasoning summaries were processed by Gemma-4-31B into more genuine reasoning like you might find on open-weight reasoning models. This won't recover the original thinking, but this should help improve training convergence when paired with other reasoning datasets. Thanks to NVIDIA NIM and Parasail. And especially thanks big thanks to https://huggingface.co/datasets/Roman1111111/claude-sonnet-4.6-120000x Note that the vast majority of the reasoning was constructed with… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/claude-sonnet-4.6-processed-reasoning.texttext-generation100K<n<1M2 likes31 downloads5mo agoHugging Face21lianghsun /tw-processed-judgments-14Bgated Dataset Card for tw-processed-judgments-14B tw-processed-judgments-14B 是一個中華民國司法判決書之大規模預訓練語料,收錄自 1996 年 1 月至最新之各審級判決書,總計約 14B tokens(約 40 GB,估計約 2,300 萬筆判決)。資料經專業法律領域知識與自然語言處理方法共同清理,移除裁定、簡易庭及格式不符之判決,僅保留具完整主文、事實、理由結構之判決書,適用於繁體中文法律領域 LLM 之預訓練與持續預訓練(CPT)。 Dataset Details Dataset Description 原始司法院判決書公開文本存在大量格式雜訊、錯誤內容與缺漏段落,導致許多以此為語料訓練之台灣本地 LLM 雖接觸過大量判決書,仍無法有效提升法律領域之推理能力。本資料集針對此問題進行大量清理與篩選: 移除裁定類型(僅保留「判決」); 移除簡易庭判決; 篩選符合標準格式之判決書(包含主文、事實、理由三大段落);… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-judgments-14B.texttext-generation1M<n<10M7 likes30 downloads6mo agoHugging Face22eewer /qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed Qwen3 4B Thinking SFT v54 Processed Training View This dataset is the processed and filtered training view used by the v54 Qwen3-4B-Thinking SFT recipe. It starts from eewer/swerebench-traces-raw-source-targeted-limitations-compaction-full-20260616-2030 and uses the strict-passed raw2030 mini-swe aligned view. Rows are compressed JSONL.zst files under data/. Each row contains a top-level messages column, optional tools, and scalar source mapping fields such as source_uuid… See the full description on the dataset page: https://huggingface.co/datasets/eewer/qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed.text-generation1K<n<10K0 likes30 downloads4mo agoHugging Face23lianghsun /tw-processed-law-ctx Dataset Card for tw-processed-law-ctx tw-processed-law-ctx 是一個中華民國(台灣)法規全文之後處理資料集,合計 11,462 筆。相較於 tw-processed-law-article(以單一條文為單位),本資料集將「同一法規之所有條文」合併為單一 text,提供整部法規之完整上下文,適合作為繁中法律 LLM 之持續預訓練(CPT)素材。 Dataset Details Dataset Description 以「條文為單位」之語料雖便於精確查詢,但模型在訓練時難以學到整部法規之章節結構與條文之間之關聯。本資料集將每一部法規之所有條文依序合併為單一長文本,保留: 法規名稱(含章節標題); 所有條文之條號與本文; 該法規最近修正日期與 API 更新日期; 若法規已廢止,額外提供 abandon_note 欄位。 每筆亦同時附帶 level(法規位階:憲法 / 法律 / 命令 / 廢止)作為篩選依據。 Curated by: Liang Hsun… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-ctx.texttext-generation10K<n<100K0 likes29 downloads6mo agoHugging Face24Fizzarolli /rpguild_processedpreprocessed version of chargoddard/rpguild into a fun little prompt format for finetuning texttext-generation10K<n<100K4 likes28 downloads2y agoHugging Face25TheFinAI /en-edgar-processedgated SEC EDGAR filings (processed) 🌐 The Fin AI Pretraining / reference corpus released by The Fin AI. Source: U.S. SEC EDGAR filings — https://www.sec.gov/search-filings. Note: Some parquet files in this repository are empty (0 bytes). Source Text extracted from public filings on the SEC EDGAR system. Structure Rows: 2,839,488 Columns: text Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/en-edgar-processed"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/en-edgar-processed.texttext-generation1M<n<10M0 likes28 downloads3d agoHugging Face26TheFinAI /gr_processed_datasetsgated Greek processed corpora 🌐 The Fin AI Processed Greek corpora (academic articles, company filings, general text, laws and regulations) as JSONL. Task pretraining corpus Language el License cc-by-4.0 text-generation0 likes24 downloads3d agoHugging Face27Process-Venue /Hindi-Marathi-Synonyms Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह) Overview This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications. The dataset provides word-synonym pairs that can be used for tasks like: Semantic analysis Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.texttext-classification1K<n<10K0 likes24 downloads2y agoHugging Face28asd-ml /roleplay_boundaries_processedProcessed version of limloop/roleplay_knowledge_boundaries. Each sample is a chat in the messages format: a system persona description (role, knowledge domains, taboo topics, communication style, evasion techniques) followed by alternating user / assistant turns. The system message also asks for short, in-character replies. messages: list of {"role", "content"} language: en or ru Train split: 3771 rows. texttext-generation1K<n<10K0 likes24 downloads2d agoHugging Face29profgabrielramos /docs-asof-processed profgabrielramos/docs-asof-processed Objetivo Dataset textual preparado para auditoria de qualidade, limpeza reprodutível e uso em pipelines de RAG. Origem Documentos publicados em profgabrielramos/docs-asof-processed. Versão v1.0.1 Estatísticas públicas (auditoria) Documentos: 803 Coluna textual: text Duplicação normalizada: 2.24% Nulos na coluna textual: 0.00% Tamanho médio (chars): 1579.1 p50/p90 (chars): 502.0 / 2221.6 Idioma… See the full description on the dataset page: https://huggingface.co/datasets/profgabrielramos/docs-asof-processed.texttext-retrievaln<1K0 likes23 downloads8mo agoHugging Face30Process-Venue /Language_Identification_v1 Dataset Card for Language Identification Dataset Dataset Summary A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications. Languages and Distribution Language Distribution: Urdu 1000 Hindi 1000 Odia 1000 Tamil 1000 Kannada 1000 Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.texttext-classification1K<n<10K1 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.