datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized
상세
데이터셋 설명
OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다.
OpenAI gpt-4o-mini를 통해 번역됐습니다.
Shared by llami-team
Language(s) (NLP): Korean
Uses
한국어 reasoning 모델 distillation
reasoning cold-start 데이터셋
Dataset Structure
question: 질문
reasoning: 추론 과정
response: 응답
Dataset Creation
[LLAMI Team] (https://llami.net)
LLAMI Github
lemon-mint
Source Data
OpenThoughts-114k-Normalized
vietnam-normalize-24kturkish-text-normalization
🇹🇷 Turkish Text Normalization (TN / ITN)
A deterministic, rule-based dataset of Turkish written ↔ spoken pairs for
Text Normalization (TN) and Inverse Text Normalization (ITN) — mapping digit/symbol
forms (1.500 TL, %25, 15.07.2026) to their fully spoken Turkish words
(bin beş yüz lira, yüzde yirmi beş, on beş temmuz iki bin yirmi altı) and back.
This is a common, high-value preprocessing step for Turkish ASR post-processing and
TTS front-ends, where numbers, dates, currencies… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-text-normalization.turkish-chat-normalization-mini
Turkish Chat Normalization Mini
turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish.
The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.normal-switch-ac76aa
normal-switch-ac76aa
Synthetic weather test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/meridianwing/normal-switch-ac76aa.basque_dialect_normalizationgeorges-1913-normalization
Normalized Georges 1913
Description
This dataset was created as part of the Burchard's Dekret Digital project (www.burchards-dekret-digital.de),
funded by the Academy of Sciences and Literature | Mainz.
It is based on 55,000 lemmata from Karl Georges, Ausführliches lateinisch-deutsches Handwörterbuch, Hannover 1913 (Georges 1913)
and was developed to train models for normalization tasks in the context of medieval Latin.
The dataset consists of approximately 5 million… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/georges-1913-normalization.carrot-engine-normalization-translation-v2config-normalization-0909Controlled path-normalization probe. No third-party data.
normalisationMnist_normalizedMnist normalized data between 0 and 1 instead of 0 and 255
include 1 - hot as label - 10 labels
every sample - 28x28 image pixel number (784)
10 labels
= 794
chinese-lexical-normalization
chinese-lexical-normalization
This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations.
Example usage:
from datasets import load_dataset
dataset = load_dataset("larrylawl/chinese-lexical-normalization")
khmer-tst-normal2royal
Khmer Text Style Transformation Dataset (Normal to Royal)
This project contains a comprehensive collection of 805 Khmer language entries, specifically designed to demonstrate the conversion of "Common/Normal" Khmer into "Royal" Khmer (រាជស័ព្ទ).
1. Content Overview
The data covers a wide variety of contexts, including:
Historical accounts: Life of King Norodom Sihanouk and historical events.
Royal Traditions: Royal ceremonies (Water Festival, Ploughing Ceremony)… See the full description on the dataset page: https://huggingface.co/datasets/k1mhor/khmer-tst-normal2royal.asturian-normalization-50k
Asturian Normalization Dataset (50k)
Dataset de 50.000 pares entrada→salida pa la normalización asturiana.
myanmar_speech_hate_and_normalclinical-narrative-implicit-normalization-bias-v0.4
Implicit Normalization Bias
Clinical Narrative Integrity v0.4
Purpose
This dataset tests whether a model:
Avoids assuming normality when data is missing
Resists default reassurance
Preserves honest narrative boundaries
Treats “normal” as a claim, not a default
You are measuring baseline discipline.
Why this dataset exists
Clinical notes often omit information.
A failure mode distinct from hallucinated negatives is more subtle:
Turning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-implicit-normalization-bias-v0.4.CodeMix_Query_Normalization_India
CodeMix Query Normalization (India)
Overview
This dataset contains code-mixed user queries from Indian contexts, primarily in Hinglish and Punjabi, normalized into clean English. It reflects how users naturally communicate in real-world scenarios by mixing local languages with English.
Features
100 high-quality samples
Code-mixed queries (Hinglish, Punjabi)
Clean normalized English outputs
Real-world, informal user language patterns
Covers domains such as… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/CodeMix_Query_Normalization_India.Advanced_CodeMix_Normalization_Dataset_India
Evaluation & Benchmarking
To validate dataset usefulness, normalization accuracy can be evaluated using:
Exact Match Accuracy
BLEU Score for text similarity
Human evaluation for real-world correctness
This dataset is designed to improve performance of multilingual NLP systems in handling noisy, code-mixed Indian queries.
Data Transformation Approach
The dataset was created by transforming real-world code-mixed queries into structured English. Variations include:… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Advanced_CodeMix_Normalization_Dataset_India.formatted_genz_normal_engTachygraphy-Microtext-Analysis-And-Normalizationarchitecture_vs_normal_image_promptsText-Normalization-Hindidiffing-stats-spp_normal10_3b-spp10-L13-Crosscoder-s1-t100-k100-lr1e-04-x32carrot-engine-normalization-translationabsa_bca_with_word_normalizationnoisy-medicine-normalizernoisy-medicine-with-lasa-normalizermedicine-normalizer-datasetdiffing-stats-spp_normal_3b-spp-L13-Crosscoder-s1-t100-k100-lr1e-04-x32_e3carrot-engine-normalization
