Team Ai
Datasetpublic

yagmurtuncer/turkish-chat-normalization-mini

Turkish Chat Normalization Mini turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish. The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes50downloads
Dataset Card

Turkish Chat Normalization Mini

turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish.

The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input side is created through controlled rule-based degradation patterns. The synthetic-data tag refers to this generated noisy-input side, not to fully artificial target text.

Dataset Summary

This dataset can be used for Turkish text-to-text experiments such as:

  • —spelling and typo correction
  • —Turkish diacritics restoration
  • —grammar and sentence cleanup
  • —informal-to-standard rewriting
  • —informal-to-formal rewriting
  • —message polishing
  • —academic or report-style polishing
  • —post-ASR text cleanup prototypes

The current release contains 20,000 examples:

SplitRows
Train16,000
Test4,000
Total20,000

Example Rows

input_textnormalized_texttask_type
Ben iyi bir ogretmen oldugumu biliyorum.Ben iyi bir öğretmen olduğumu biliyorum.diacritics_restoration
yagmur yagarsa ve toprak nemli olursa, otarı cekmek daha kolayy olurYağmur yağarsa ve toprak nemli olursa, otları çekmek daha kolay olur.typo_correction
biz ayrıldık birbirimizi asla tekrar gormeyecegiz.Biz ayrıldık, birbirimizi asla tekrar görmeyeceğiz.grammar_fix
sizin begenmediginiz bi seyi baskası begense bile size fazla bi yararı olmaz.Sizin beğenmediğiniz bir şeyi başkası beğense bile size fazla bir yararı olmaz.informal_to_standard
hocam yeni bi palet ve birkac boyama fırcası aldım.Yeni bir palet ve birkaç boyama fırçası aldım.informal_to_formal
ornegin 2+3*7 gibi bir ifadede once hangi islem yapılacak?Örneğin 2+3*7 gibi bir ifadede önce hangi işlem yapılacak?message_polishing
bellekteki ardısık gozeneklerin toplamına dizi denir.Bellekteki ardışık gözeneklerin toplamına dizi denir.grammar_fix
ayrıca ciftcilerimize hububatta ton basına 230 lira prim destek odemesi yapıyoruz.Ayrıca, çiftçilerimize hububatta ton başına 230 lira prim ve destek ödemesi yapıyoruz.academic_polishing

Data Sources

The dataset is derived from multiple open Turkish text sources:

SourceRowsShareLicense note
Tatoeba Turkish sentence export7,03335.16%CC BY 2.0 FR
Turkish Wikibooks3,27016.35%CC BY-SA 4.0
Turkish Wikisource3,26216.31%CC BY-SA 4.0
Turkish Wikipedia3,24516.23%CC BY-SA 4.0
Turkish Wikiquote3,19015.95%CC BY-SA 4.0

Each row includes source and license fields so that downstream users can inspect provenance and preserve attribution.

Schema

The CSV files contain the following columns:

ColumnDescription
idRow identifier within the split
input_textNoisy, incomplete, informal, or degraded input text
normalized_textClean target text
task_typeType of normalization task
styleExpected target style, such as casual, standard, formal, or academic
domainApproximate topic label inferred from source keywords
sourceSource collection name
licenseSource license note
splittrain or test

Task Types

TaskDescriptionRowsShare
typo_correctionCorrects misspellings, dropped letters, or character swaps2,85714.29%
diacritics_restorationRestores Turkish characters such as ç, ğ, ı, ö, ş, and ü2,85714.29%
grammar_fixProduces a cleaner and more grammatical sentence2,85614.28%
informal_to_standardRewrites informal text into standard Turkish2,85714.29%
informal_to_formalRewrites informal text with a more formal tone2,85914.30%
message_polishingImproves message clarity, tone, and readability2,85814.29%
academic_polishingMoves the text closer to academic or report-style Turkish2,85614.28%

Style and Domain Labels

The style and domain fields are not required for basic supervised training, but they are useful for filtering, evaluation, error analysis, and controlled experiments.

Style Distribution

StyleDescriptionRowsShare
casualEveryday or low-formality language5,71428.57%
standardNeutral standard Turkish5,71328.56%
formalMore formal and structured wording5,71728.58%
academicAcademic or report-style wording2,85614.28%

Domain Distribution

DomainDescriptionRowsShare
formal_public_textGeneral encyclopedic, public, or neutral text13,98869.94%
technical_supportTechnical, software, data, system, or internet-related text4,57622.88%
product_reviewProduct, order, delivery, price, or shopping-related text7593.80%
student_messageEducation, course, school, university, or academic context6773.38%

Files

  • —data/train.csv: training split
  • —data/test.csv: test split
  • —scripts/generate_dataset.py: source collection and controlled degradation pipeline
  • —scripts/pii_filter.py: scans CSV files for common PII-like patterns
  • —scripts/validate_dataset.py: validates schema, splits, required fields, and distribution summaries

Quality Checks

The released CSV files were checked for:

  • —consistent column names across train and test files
  • —UTF-8 CSV readability
  • —non-empty input_text and normalized_text fields
  • —valid train / test split labels
  • —expected 16,000 / 4,000 split sizes
  • —common PII-like patterns such as emails, phone numbers, URL-like strings, and Turkish-ID-like numeric patterns
  • —basic row-level integrity, including duplicate IDs within each split

Generation Method

The dataset was built with a deterministic normalization-pair generation pipeline:

  1. 1.Turkish source sentences were collected from open, attribution-friendly web resources.
  2. 2.Source sentences were cleaned and retained as normalized_text.
  3. 3.Controlled degradation rules were applied to create input_text, including diacritics removal, lowercasing, typo injection, punctuation simplification, connector removal, and informal abbreviation replacement.
  4. 4.Each example was labeled with task_type, style, domain, source, license, and split.
  5. 5.The final dataset was split into train and test subsets using an 80/20 ratio.

Intended Use

This dataset is intended for small and medium-scale Turkish text-to-text experiments, including:

  • —sequence-to-sequence normalization models
  • —LLM instruction-tuning prototypes
  • —Turkish spelling and style correction pipelines
  • —post-ASR text normalization experiments
  • —evaluation of Turkish rewriting and cleanup systems

Example:

text
Input: hocam bugun toplantıya gelemicem raporu yarin aticam
Output: Bugün toplantıya gelemeyeceğim. Raporu yarın ileteceğim.

Limitations

This dataset is designed as a controlled benchmark for Turkish text normalization. The noisy inputs were generated from openly licensed Turkish text sources using rule-based degradation patterns. As a result, the dataset is suitable for typo correction, diacritics restoration, grammar cleanup, and style-transfer experiments, but it may not fully represent all informal writing styles found in real-world conversations.

The dataset is not intended for toxicity detection, privacy classification, or safety moderation. Downstream users should preserve source attribution and review license compatibility for their use case.

Citation and Attribution

If you use this dataset, please cite or mention this dataset name and preserve attribution to the source projects:

  • —Tatoeba Turkish sentence export
  • —Turkish Wikipedia
  • —Turkish Wikibooks
  • —Turkish Wikiquote
  • —Turkish Wikisource

Notes

This dataset intentionally avoids Reddit, YouTube comments, X/Twitter, Instagram, TikTok, complaint platforms, private chat logs, and other user-generated conversational data sources.