Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes180 downloads2y agoHugging Face02thanhkt /vietnam-normalize-24ktexttext-generation10K<n<100K3 likes61 downloads2y agoHugging Face03yagmurtuncer /turkish-text-normalization 🇹🇷 Turkish Text Normalization (TN / ITN) A deterministic, rule-based dataset of Turkish written ↔ spoken pairs for Text Normalization (TN) and Inverse Text Normalization (ITN) — mapping digit/symbol forms (1.500 TL, %25, 15.07.2026) to their fully spoken Turkish words (bin beş yüz lira, yüzde yirmi beş, on beş temmuz iki bin yirmi altı) and back. This is a common, high-value preprocessing step for Turkish ASR post-processing and TTS front-ends, where numbers, dates, currencies… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-text-normalization.texttext-generation10K<n<100K0 likes61 downloads3mo agoHugging Face04yagmurtuncer /turkish-chat-normalization-mini Turkish Chat Normalization Mini turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish. The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.texttext-generation10K<n<100K0 likes50 downloads4mo agoHugging Face05meridianwing /normal-switch-ac76aa normal-switch-ac76aa Synthetic weather test data: 46 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/meridianwing/normal-switch-ac76aa.tabularn<1K0 likes47 downloads1mo agoHugging Face06jaio98 /basque_dialect_normalizationtext1K<n<10K0 likes35 downloads4mo agoHugging Face07mschonhardt /georges-1913-normalization Normalized Georges 1913 Description This dataset was created as part of the Burchard's Dekret Digital project (www.burchards-dekret-digital.de), funded by the Academy of Sciences and Literature | Mainz. It is based on 55,000 lemmata from Karl Georges, Ausführliches lateinisch-deutsches Handwörterbuch, Hannover 1913 (Georges 1913) and was developed to train models for normalization tasks in the context of medieval Latin. The dataset consists of approximately 5 million… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/georges-1913-normalization.text1M<n<10M0 likes31 downloads2y agoHugging Face08sellersew /carrot-engine-normalization-translation-v2text10M<n<100M1 likes23 downloads3y agoHugging Face09authormist /config-normalization-0909Controlled path-normalization probe. No third-party data. textn<1K0 likes22 downloads1mo agoHugging Face10Alexandre-Numind /normalisationtextn<1K0 likes21 downloads2y agoHugging Face11larrylawl /chinese-lexical-normalization chinese-lexical-normalization This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations. Example usage: from datasets import load_dataset dataset = load_dataset("larrylawl/chinese-lexical-normalization") text1K<n<10K0 likes18 downloads3y agoHugging Face12k1mhor /khmer-tst-normal2royal Khmer Text Style Transformation Dataset (Normal to Royal) This project contains a comprehensive collection of 805 Khmer language entries, specifically designed to demonstrate the conversion of "Common/Normal" Khmer into "Royal" Khmer (រាជស័ព្ទ). 1. Content Overview The data covers a wide variety of contexts, including: Historical accounts: Life of King Norodom Sihanouk and historical events. Royal Traditions: Royal ceremonies (Water Festival, Ploughing Ceremony)… See the full description on the dataset page: https://huggingface.co/datasets/k1mhor/khmer-tst-normal2royal.texttext-generationn<1K0 likes18 downloads10mo agoHugging Face13ajescandon /asturian-normalization-50k Asturian Normalization Dataset (50k) Dataset de 50.000 pares entrada→salida pa la normalización asturiana. text10K<n<100K0 likes8 downloads11mo agoHugging Face14Gojokun00 /myanmar_speech_hate_and_normaltext1K<n<10K1 likes8 downloads8mo agoHugging Face15ClarusC64 /clinical-narrative-implicit-normalization-bias-v0.4 Implicit Normalization Bias Clinical Narrative Integrity v0.4 Purpose This dataset tests whether a model: Avoids assuming normality when data is missing Resists default reassurance Preserves honest narrative boundaries Treats “normal” as a claim, not a default You are measuring baseline discipline. Why this dataset exists Clinical notes often omit information. A failure mode distinct from hallucinated negatives is more subtle: Turning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-implicit-normalization-bias-v0.4.tabularn<1K0 likes6 downloads9mo agoHugging Face16karanverma19 /CodeMix_Query_Normalization_India CodeMix Query Normalization (India) Overview This dataset contains code-mixed user queries from Indian contexts, primarily in Hinglish and Punjabi, normalized into clean English. It reflects how users naturally communicate in real-world scenarios by mixing local languages with English. Features 100 high-quality samples Code-mixed queries (Hinglish, Punjabi) Clean normalized English outputs Real-world, informal user language patterns Covers domains such as… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/CodeMix_Query_Normalization_India.textn<1K0 likes6 downloads6mo agoHugging Face17karanverma19 /Advanced_CodeMix_Normalization_Dataset_India Evaluation & Benchmarking To validate dataset usefulness, normalization accuracy can be evaluated using: Exact Match Accuracy BLEU Score for text similarity Human evaluation for real-world correctness This dataset is designed to improve performance of multilingual NLP systems in handling noisy, code-mixed Indian queries. Data Transformation Approach The dataset was created by transforming real-world code-mixed queries into structured English. Variations include:… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Advanced_CodeMix_Normalization_Dataset_India.textn<1K0 likes6 downloads6mo agoHugging Face18jkb2002 /formatted_genz_normal_engtextn<1K0 likes5 downloads3y agoHugging Face19archismancoder /Tachygraphy-Microtext-Analysis-And-Normalizationgatedtext10K<n<100K0 likes5 downloads2y agoHugging Face20as-cle-bert /architecture_vs_normal_image_promptstext1K<n<10K2 likes5 downloads2y agoHugging Face21shubham-Bgs /Text-Normalization-Hinditext1K<n<10K1 likes5 downloads2y agoHugging Face22Sitavi /diffing-stats-spp_normal10_3b-spp10-L13-Crosscoder-s1-t100-k100-lr1e-04-x32tabular10K<n<100K0 likes5 downloads2mo agoHugging Face23riggj /carrot-engine-normalization-translationtextn<1K0 likes4 downloads3y agoHugging Face24mahfuzh74 /absa_bca_with_word_normalizationtext100K<n<1M0 likes4 downloads2y agoHugging Face25AshwinManohar /noisy-medicine-normalizertext10K<n<100K0 likes4 downloads1y agoHugging Face26AshwinManohar /noisy-medicine-with-lasa-normalizertext10K<n<100K0 likes4 downloads1y agoHugging Face27AshwinManohar /medicine-normalizer-datasettext1K<n<10K0 likes4 downloads1y agoHugging Face28Sitavi /diffing-stats-spp_normal_3b-spp-L13-Crosscoder-s1-t100-k100-lr1e-04-x32_e3tabular10K<n<100K0 likes4 downloads3mo agoHugging Face29riggj /carrot-engine-normalizationtextn<1K0 likes3 downloads3y agoHugging Face30thucnc /address_normalizertextn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.