Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.1k downloads4mo agoHugging Face02vllm-sr /halueval-spans-normalized HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts) 🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans-normalized") Why Normalized Prompts? Training on mixed datasets with different prompt formats causes distribution shift: Original Format… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans-normalized.texttoken-classification10K<n<100K0 likes41 downloads9mo agoHugging Face03Kicshikxo /dataset-context-self.qwen2.5-max2048.normalized-lower.nohello-v5.max384text1M<n<10M0 likes32 downloads14d agoHugging Face04lemon-mint /OpenThoughts-114k-Normalizedprefixes = [ "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.", "Return your final response within \\boxed{}. ", "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.", ] // -1 if None text100K<n<1M1 likes30 downloads2y agoHugging Face05Jnx03 /kanitakorn-deepseek-v44-normalized-mcq-replay-mix Kanitakorn v44 normalized MCQ replay mix Original and previously audited synthetic SFT mixture for a non-Thai-base DeepSeek/Qwen-style <=14B candidate. This dataset does not include benchmark prompts, benchmark gold answers, model benchmark samples, BoN traces, routing labels, or consensus outputs. Design intent: normalize MCQ final marker to คำตอบคือ (x) keep explanations before the answer instead of answer-only rows target aggregate ThaiExam failure buckets: grammar/spelling… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v44-normalized-mcq-replay-mix.text1K<n<10K0 likes28 downloads4mo agoHugging Face06cometadata /2025-08-datacite-normalized-affiliation-string-distribution DataCite Normalized Affiliation Distribution Summary normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export. Structure { "normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.text1M<n<10M0 likes26 downloads11mo agoHugging Face07deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes25 downloads5mo agoHugging Face08francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face09PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_both_normalizedtext1K<n<10K0 likes15 downloads2y agoHugging Face10atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face11PhdDz /PubmedQA_5_WITH_RELATION_vqc_primekg_normalizedtext1K<n<10K0 likes13 downloads2y agoHugging Face12joduor /adaption-gd-unk-normalized-samples This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-gd_unk_normalized_samples This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement. Dataset size There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.textn<1K0 likes13 downloads5mo agoHugging Face13PhdDz /PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_hetionet_normalizedtext1K<n<10K0 likes11 downloads2y agoHugging Face14PhdDz /PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes11 downloads2y agoHugging Face15PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes11 downloads2y agoHugging Face16EsotericsEnjoyer /CPT-normalizedReplaced special characters where it makes a difference: what was: 10  2 ­ 16 ​ 19 § 2 © 327 « 10 ¬ 1 ® 85 ° 3 ´ 16 ¶ 323 » 31 × 13 ÷ 2 ˚ 1 ‑ 10683 – 88952 — 2409 ‘ 5967 ’ 4663 “ 4223 ” 9 † 81 • 355 … 18 ′ 19 ‹ 19 › 12 → 15 − 1 √ 5 ∞ 101 ∴ 2 ≥ 2 ⊣ 1 □ 5 △ 4 ▽ 1 ◒ 5 ☉ 2 ☋ 1 ☥ 3 ☽… See the full description on the dataset page: https://huggingface.co/datasets/EsotericsEnjoyer/CPT-normalized.text1K<n<10K0 likes11 downloads8mo agoHugging Face17PhdDz /PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_both_normalizedtext1K<n<10K0 likes9 downloads2y agoHugging Face18PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_hetionet_normalizedtext1K<n<10K0 likes9 downloads2y agoHugging Face19PhdDz /PubmedQA_5_WITH_DEFINITION_vsimilarity_both_normalizedtext1K<n<10K0 likes8 downloads2y agoHugging Face20PhdDz /PubmedQA_5_WITH_DEFINITION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes8 downloads2y agoHugging Face21morning-light /gpt-oss20b-samples_normalized 关于数据集 数据集来源:gpt-oss20b-samples_deduplicated 数据集清洗流程: 使用ASCII以及fasttext判断乱码; 对数据集内标记语言进行删除(如Markdown、HTML等); 使用SimHash进行去重 多次前缀去重; 清洗字符数小于2000的数据 初始数据(转jsonl后): 样本总数: 183063; 总字符数: 4408366009; 平均长度: 24081.14 字符 乱码清洗和去除标记语言后: 样本总数: 177582; 总字符数: 3463945036; 平均长度: 19506.17 字符 相似度去重、去前缀后(十分之一多点的样本数据): 样本总数: 17859; 总字符数: 189850733; 平均长度: 10630.54 字符 最后处理完的全部数据… See the full description on the dataset page: https://huggingface.co/datasets/morning-light/gpt-oss20b-samples_normalized.text100K<n<1M0 likes8 downloads10mo agoHugging Face22PhdDz /PubmedQA_5_WITH_DEFINITION_vsimilarity_hetionet_normalizedtext1K<n<10K0 likes7 downloads2y agoHugging Face23PhdDz /BioASQBlurb_5_WITH_DEFINITION_RELATION_vsimilarity_both_normalizedtextn<1K0 likes6 downloads2y agoHugging Face24PhdDz /BioASQBlurb_5_WITH_RELATION_vsimilarity_both_normalizedtextn<1K0 likes6 downloads2y agoHugging Face25PhdDz /BioASQBlurb_5_WITH_RELATION_vsimilarity_hetionet_normalizedtextn<1K0 likes6 downloads2y agoHugging Face26PhdDz /BioASQBlurb_5_WITH_RELATION_vsimilarity_primekg_normalizedtextn<1K0 likes6 downloads2y agoHugging Face27PhdDz /BioASQ_5_WITH_RELATION_vsimilarity_both_normalizedtext1K<n<10K0 likes6 downloads2y agoHugging Face28PhdDz /PubmedQA_5_WITH_RELATION_vqc_both_normalizedtext1K<n<10K0 likes6 downloads2y agoHugging Face29aemmeath /Aether_reasoning_normalized_v2text1M<n<10M0 likes6 downloads4mo agoHugging Face30PhdDz /BioASQBlurb_5_WITH_DEFINITION_RELATION_vsimilarity_hetionet_normalizedtextn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.