Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.1k downloads4mo agoHugging Face02Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes156 downloads1mo agoHugging Face03sraivante /Financial-Form-Normalization-Instructions Financial Form Normalization Instructions 3,657 English instruction examples across 38 tasks, associated with sraivante/TinyLlama-1.1B-Financial-Form-Normalizer-LoRA. The examples teach short user replies to map to predefined application field values: dates, amounts, ZIP codes, yes/no or boolean values, category labels, navigation intents and small stage/state JSON objects. The intended task is supplied by a system prompt. Copyright (c) 2026 sraivante, for original dataset… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/Financial-Form-Normalization-Instructions.texttext-generation1K<n<10K0 likes79 downloads16d agoHugging Face04zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes64 downloads1mo agoHugging Face05vllm-sr /halueval-spans-normalized HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts) 🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans-normalized") Why Normalized Prompts? Training on mixed datasets with different prompt formats causes distribution shift: Original Format… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans-normalized.texttoken-classification10K<n<100K0 likes41 downloads9mo agoHugging Face06skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes39 downloads2mo agoHugging Face07Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes33 downloads5mo agoHugging Face08adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes33 downloads1mo agoHugging Face09Kicshikxo /dataset-context-self.qwen2.5-max2048.normalized-lower.nohello-v5.max384text1M<n<10M0 likes32 downloads14d agoHugging Face10lemon-mint /OpenThoughts-114k-Normalizedprefixes = [ "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.", "Return your final response within \\boxed{}. ", "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.", ] // -1 if None text100K<n<1M1 likes30 downloads2y agoHugging Face11lunahr /normalization-data-mixed Normalization Dataset (Mixed) This dataset is a collection of 50000 rows originating from various sources: Wikipedia - 20000 rows PersonaChat truecased - 20000 rows Synthetic edge case data - 5000 rows Synthetic quoted text data - 5000 rows The synthetic data has been generated using GPT-5.3 models. The other data was sourced from the original Hugging Face sources. This dataset can be used to train text normalizers that convert badly formatted English into correct English.… See the full description on the dataset page: https://huggingface.co/datasets/lunahr/normalization-data-mixed.text10K<n<100K0 likes29 downloads3mo agoHugging Face12Jnx03 /kanitakorn-deepseek-v44-normalized-mcq-replay-mix Kanitakorn v44 normalized MCQ replay mix Original and previously audited synthetic SFT mixture for a non-Thai-base DeepSeek/Qwen-style <=14B candidate. This dataset does not include benchmark prompts, benchmark gold answers, model benchmark samples, BoN traces, routing labels, or consensus outputs. Design intent: normalize MCQ final marker to คำตอบคือ (x) keep explanations before the answer instead of answer-only rows target aggregate ThaiExam failure buckets: grammar/spelling… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v44-normalized-mcq-replay-mix.text1K<n<10K0 likes28 downloads4mo agoHugging Face13cometadata /2025-08-datacite-normalized-affiliation-string-distribution DataCite Normalized Affiliation Distribution Summary normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export. Structure { "normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.text1M<n<10M0 likes26 downloads11mo agoHugging Face14deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes25 downloads5mo agoHugging Face15francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face16open-llm-leaderboard /marcuscedricridia__cursa-o1-7b-v1.2-normalize-false-detailsgated Dataset Card for Evaluation run of marcuscedricridia/cursa-o1-7b-v1.2-normalize-false Dataset automatically created during the evaluation run of model marcuscedricridia/cursa-o1-7b-v1.2-normalize-false The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__cursa-o1-7b-v1.2-normalize-false-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face17open-llm-leaderboard /ehristoforu__fq2.5-7b-it-normalize_true-detailsgated Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_true Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_true The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_true-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face18ruscorpora /normalization TL;DR: Text Normalization for Social Media Corpus Dataset Description This dataset contains examples of Russian-language texts from social networks with distorted spelling (typos, abbreviations, etc.) and their normalized versions in json format. A detailed spelling correction protocol is given in the TBA article. The dataset size is 1930 sentence pairs. In each pair, the sentences are tokenized by words, and the lengths of both sentences in the pair are equal. If a… See the full description on the dataset page: https://huggingface.co/datasets/ruscorpora/normalization.text1K<n<10K0 likes18 downloads1y agoHugging Face19open-llm-leaderboard /ehristoforu__fq2.5-7b-it-normalize_false-detailsgated Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_false Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_false The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_false-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face20gabormarko /franka-insert-siemens-lid-eval-normal-50imagen<1K0 likes17 downloads5mo agoHugging Face21PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_both_normalizedtext1K<n<10K0 likes15 downloads2y agoHugging Face22happy8825 /after_incident_normaltextn<1K0 likes15 downloads10mo agoHugging Face23atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face24normalcomputing /wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA Github: https://github.com/normal-computing/extended-mind-transformers/ ArXiv: https://arxiv.org/abs/2406.02332 Original dataset by Abacus AI. Developed by: Normal Computing, Adapted from Abacus AI License: Apache 2.0 Long-range Counterfactual Retrieval Benchmark This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.textn<1K1 likes13 downloads2y agoHugging Face25PhdDz /PubmedQA_5_WITH_RELATION_vqc_primekg_normalizedtext1K<n<10K0 likes13 downloads2y agoHugging Face26happy8825 /normal_ecva_sfttextn<1K0 likes13 downloads10mo agoHugging Face27joduor /adaption-gd-unk-normalized-samples This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-gd_unk_normalized_samples This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement. Dataset size There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.textn<1K0 likes13 downloads5mo agoHugging Face28brendan-gho /qwen1.5b_normal_numstext10K<n<100K0 likes12 downloads5mo agoHugging Face29backup-dev /normalize_symlinkstabularn<1K0 likes12 downloads5mo agoHugging Face30PhdDz /PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_hetionet_normalizedtext1K<n<10K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.