Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SaakethS /fishnet-lichess-normalized1 likes4.2k downloads23d agoHugging Face02mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.2k downloads4mo agoHugging Face03winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes712 downloads1y agoHugging Face04jcy20 /DiffusionPDE-normalizedWe take the dataset from DiffusionPDE. For convenience, we provide our processed version on Hugging Face (see Appendix D & E in our FunDPS paper for details). The processing scripts are provided under utils/ here. It is worth noting that we normalized the datasets to zero mean and 0.5 standard deviation to follow EDM's practice. 10K<n<100K2 likes628 downloads1y agoHugging Face05TongheZhangTH /XDof-TshirtFolding-20hours-normalizedtabular1M<n<10M0 likes623 downloads9mo agoHugging Face06cfy2yue /perturbseq_normalized CellClip public normalized perturbation single-cell release This public repository contains the source-cleared, human, cell-level portion of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and 260,859,318,437 payload bytes. These are processed derivatives rather than the original raw download archives. Source partition Units Cells Reference cells Effect cells scPerturb (23 genepert + 229 chempert) 252 4,774,802 397,528 4,377,274 XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.feature-extraction0 likes551 downloads3mo agoHugging Face07N03N9 /cv24-tr-128-normalizedtext100K<n<1M0 likes353 downloads10mo agoHugging Face08winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes333 downloads1y agoHugging Face09jxie /epsilon-normalized Dataset Card for "epsilon-normalized" More Information needed 100K<n<1M0 likes259 downloads3y agoHugging Face10kgnlp /meld-open-normalized MELD Open (Normalized) MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details. Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.tabulartoken-classification10M<n<100M0 likes256 downloads5mo agoHugging Face11csoai /gspc-normalized GSPC normalised — every bank in one schema The one schema to read first. 518 rows that flatten several GSPC banks into a single shape: source (the bank repository the row came from), axis, category, anchor, prompt, expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently dropped. If you want to reuse the banks without learning each one's native layout, start here. The live board is the authority GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.textothern<1K0 likes256 downloads2d agoHugging Face12mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes255 downloads11mo agoHugging Face13Scicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes232 downloads6mo agoHugging Face14N03N9 /cv24-ur-128-normalizedtext10K<n<100K0 likes210 downloads10mo agoHugging Face15Twelve2five /igbo_tts_normalizedaudio100K<n<1M2 likes206 downloads1y agoHugging Face16N03N9 /cv24-cy-128-normalizedtext10K<n<100K0 likes194 downloads10mo agoHugging Face17TongheZhangTH /CartonPickNPlace2Target-normalizedtabular10K<n<100K0 likes189 downloads9mo agoHugging Face18N03N9 /cv24-de-128-normalizedtext100K<n<1M0 likes184 downloads10mo agoHugging Face19N03N9 /cv24-uk-128-normalizedtext10K<n<100K0 likes184 downloads10mo agoHugging Face20N03N9 /cv24-pt-128-normalizedtext100K<n<1M0 likes179 downloads10mo agoHugging Face21N03N9 /cv24-ca-128-normalizedtext1M<n<10M0 likes178 downloads10mo agoHugging Face22N03N9 /cv24-ru-128-normalizedtext100K<n<1M0 likes177 downloads10mo agoHugging Face23llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes175 downloads2y agoHugging Face24N03N9 /cv24-sw-128-normalizedtext100K<n<1M0 likes172 downloads10mo agoHugging Face25jxie /qg-tagging-normalized Dataset Card for "qg-tagging-normalized" More Information needed 1M<n<10M0 likes156 downloads3y agoHugging Face26N03N9 /cv24-sk-128-normalizedtext10K<n<100K0 likes146 downloads10mo agoHugging Face27warmestman /common-voice-20-mn-normalized Common Voice 20.0 Mongolian Dataset This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release. Dataset Structure The dataset contains: Audio clips in .mp3 format Transcriptions for each audio clip Train/test/dev splits Additional metadata including speaker demographics Usage This dataset can be used for: Speech Recognition Voice Analysis Linguistic Research Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.audioautomatic-speech-recognition10K<n<100K4 likes141 downloads2y agoHugging Face28Eimhin03 /Fleurs_Irish_normalizedaudio1K<n<10K0 likes129 downloads6mo agoHugging Face29MoaazTalab /ASVspoof_2021_DF_Balanced_Normalizedaudio100K<n<1M5 likes128 downloads2y agoHugging Face30vc940 /business-entity-resolution-normalized Business Entity Resolution: normalised records Normalised copies of the six source files of the ML Challenge 2026 Business Entity Resolution task (business records from three sources, US / India in train, plus France in test). The goal of the task is to find, for every Source 1 record, the Source 2 / Source 3 records that describe the same business. File Rows train_s1.parquet 2,206,821 train_s2.parquet 5,034,616 train_s3.parquet 5,285,603 test_s1.parquet 1,732… See the full description on the dataset page: https://huggingface.co/datasets/vc940/business-entity-resolution-normalized.tabular10M<n<100M0 likes112 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.