datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.halueval-spans-normalized
HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts)
🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility.
Quick Start
from datasets import load_dataset
dataset = load_dataset("llm-semantic-router/halueval-spans-normalized")
Why Normalized Prompts?
Training on mixed datasets with different prompt formats causes distribution shift:
Original Format… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans-normalized.dataset-context-self.qwen2.5-max2048.normalized-lower.nohello-v5.max384OpenThoughts-114k-Normalizedprefixes = [
"Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.",
"Return your final response within \\boxed{}. ",
"Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.",
]
// -1 if None
kanitakorn-deepseek-v44-normalized-mcq-replay-mix
Kanitakorn v44 normalized MCQ replay mix
Original and previously audited synthetic SFT mixture for a non-Thai-base DeepSeek/Qwen-style <=14B candidate.
This dataset does not include benchmark prompts, benchmark gold answers, model benchmark samples, BoN traces, routing labels, or consensus outputs.
Design intent:
normalize MCQ final marker to คำตอบคือ (x)
keep explanations before the answer instead of answer-only rows
target aggregate ThaiExam failure buckets: grammar/spelling… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v44-normalized-mcq-replay-mix.2025-08-datacite-normalized-affiliation-string-distribution
DataCite Normalized Affiliation Distribution
Summary
normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total
occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the
provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export.
Structure
{
"normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.PubmedQA_5_WITH_RELATION_vsimilarity_both_normalizeddair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.PubmedQA_5_WITH_RELATION_vqc_primekg_normalizedadaption-gd-unk-normalized-samples
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-gd_unk_normalized_samples
This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement.
Dataset size
There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_hetionet_normalizedPubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_primekg_normalizedPubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizedCPT-normalizedReplaced special characters where it makes a difference:
what was:
10
2
16
19 §
2 ©
327 «
10 ¬
1 ®
85 °
3 ´
16 ¶
323 »
31 ×
13 ÷
2 ˚
1 ‑
10683 –
88952 —
2409 ‘
5967 ’
4663 “
4223 ”
9 †
81 •
355 …
18 ′
19 ‹
19 ›
12 →
15 −
1 √
5 ∞
101 ∴
2 ≥
2 ⊣
1 □
5 △
4 ▽
1 ◒
5 ☉
2 ☋
1 ☥
3 ☽… See the full description on the dataset page: https://huggingface.co/datasets/EsotericsEnjoyer/CPT-normalized.PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_both_normalizedPubmedQA_5_WITH_RELATION_vsimilarity_hetionet_normalizedPubmedQA_5_WITH_DEFINITION_vsimilarity_both_normalizedPubmedQA_5_WITH_DEFINITION_vsimilarity_primekg_normalizedgpt-oss20b-samples_normalized
关于数据集
数据集来源:gpt-oss20b-samples_deduplicated
数据集清洗流程:
使用ASCII以及fasttext判断乱码;
对数据集内标记语言进行删除(如Markdown、HTML等);
使用SimHash进行去重
多次前缀去重;
清洗字符数小于2000的数据
初始数据(转jsonl后):
样本总数: 183063;
总字符数: 4408366009;
平均长度: 24081.14 字符
乱码清洗和去除标记语言后:
样本总数: 177582;
总字符数: 3463945036;
平均长度: 19506.17 字符
相似度去重、去前缀后(十分之一多点的样本数据):
样本总数: 17859;
总字符数: 189850733;
平均长度: 10630.54 字符
最后处理完的全部数据… See the full description on the dataset page: https://huggingface.co/datasets/morning-light/gpt-oss20b-samples_normalized.PubmedQA_5_WITH_DEFINITION_vsimilarity_hetionet_normalizedBioASQBlurb_5_WITH_DEFINITION_RELATION_vsimilarity_both_normalizedBioASQBlurb_5_WITH_RELATION_vsimilarity_both_normalizedBioASQBlurb_5_WITH_RELATION_vsimilarity_hetionet_normalizedBioASQBlurb_5_WITH_RELATION_vsimilarity_primekg_normalizedBioASQ_5_WITH_RELATION_vsimilarity_both_normalizedPubmedQA_5_WITH_RELATION_vqc_both_normalizedAether_reasoning_normalized_v2BioASQBlurb_5_WITH_DEFINITION_RELATION_vsimilarity_hetionet_normalized
