yagmurtuncer/turkish-chat-normalization-mini
Turkish Chat Normalization Mini turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish. The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.
Turkish Chat Normalization Mini
turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish.
The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input side is created through controlled rule-based degradation patterns. The synthetic-data tag refers to this generated noisy-input side, not to fully artificial target text.
Dataset Summary
This dataset can be used for Turkish text-to-text experiments such as:
- spelling and typo correction
- Turkish diacritics restoration
- grammar and sentence cleanup
- informal-to-standard rewriting
- informal-to-formal rewriting
- message polishing
- academic or report-style polishing
- post-ASR text cleanup prototypes
The current release contains 20,000 examples:
Example Rows
Data Sources
The dataset is derived from multiple open Turkish text sources:
Each row includes source and license fields so that downstream users can inspect provenance and preserve attribution.
Schema
The CSV files contain the following columns:
Task Types
Style and Domain Labels
The style and domain fields are not required for basic supervised training, but they are useful for filtering, evaluation, error analysis, and controlled experiments.
Style Distribution
Domain Distribution
Files
data/train.csv: training splitdata/test.csv: test splitscripts/generate_dataset.py: source collection and controlled degradation pipelinescripts/pii_filter.py: scans CSV files for common PII-like patternsscripts/validate_dataset.py: validates schema, splits, required fields, and distribution summaries
Quality Checks
The released CSV files were checked for:
- consistent column names across train and test files
- UTF-8 CSV readability
- non-empty
input_textandnormalized_textfields - valid
train/testsplit labels - expected 16,000 / 4,000 split sizes
- common PII-like patterns such as emails, phone numbers, URL-like strings, and Turkish-ID-like numeric patterns
- basic row-level integrity, including duplicate IDs within each split
Generation Method
The dataset was built with a deterministic normalization-pair generation pipeline:
- Turkish source sentences were collected from open, attribution-friendly web resources.
- Source sentences were cleaned and retained as
normalized_text. - Controlled degradation rules were applied to create
input_text, including diacritics removal, lowercasing, typo injection, punctuation simplification, connector removal, and informal abbreviation replacement. - Each example was labeled with
task_type,style,domain,source,license, andsplit. - The final dataset was split into train and test subsets using an 80/20 ratio.
Intended Use
This dataset is intended for small and medium-scale Turkish text-to-text experiments, including:
- sequence-to-sequence normalization models
- LLM instruction-tuning prototypes
- Turkish spelling and style correction pipelines
- post-ASR text normalization experiments
- evaluation of Turkish rewriting and cleanup systems
Example:
Input: hocam bugun toplantıya gelemicem raporu yarin aticam
Output: Bugün toplantıya gelemeyeceğim. Raporu yarın ileteceğim.Limitations
This dataset is designed as a controlled benchmark for Turkish text normalization. The noisy inputs were generated from openly licensed Turkish text sources using rule-based degradation patterns. As a result, the dataset is suitable for typo correction, diacritics restoration, grammar cleanup, and style-transfer experiments, but it may not fully represent all informal writing styles found in real-world conversations.
The dataset is not intended for toxicity detection, privacy classification, or safety moderation. Downstream users should preserve source attribution and review license compatibility for their use case.
Citation and Attribution
If you use this dataset, please cite or mention this dataset name and preserve attribution to the source projects:
- Tatoeba Turkish sentence export
- Turkish Wikipedia
- Turkish Wikibooks
- Turkish Wikiquote
- Turkish Wikisource
Notes
This dataset intentionally avoids Reddit, YouTube comments, X/Twitter, Instagram, TikTok, complaint platforms, private chat logs, and other user-generated conversational data sources.
