Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.1k downloads4mo agoHugging Face02arijitghosh /T2I-ImageNet-Normalimage1M<n<10M3 likes1.3k downloads1y agoHugging Face03izumi-lab /mc4-ja-filter-ja-normal Dataset Card for "mc4-ja-filter-ja-normal" More Information needed text10M<n<100M5 likes819 downloads3y agoHugging Face04Pointcept /modelnet40_normal_resampled-compressedtext10K<n<100K2 likes798 downloads2y agoHugging Face05winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes733 downloads1y agoHugging Face06izumi-lab /oscar2301-ja-filter-ja-normal Dataset Card for "oscar2301-ja-filter-ja-normal" More Information needed text10M<n<100M6 likes544 downloads3y agoHugging Face07ducido /merged_libero_s100_mdnl_original_s100_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 1676, "total_frames": 269918, "total_tasks": 40, "total_videos": 0, "total_chunks": 2, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:1676" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s100_mdnl_original_s100_synthetic_20ep_2_hard_task_5ep_all_normal_task.imagerobotics100K<n<1M0 likes398 downloads4mo agoHugging Face08ducido /merged_libero_s100_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task_PAThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 1676, "total_frames": 269918, "total_tasks": 40, "total_videos": 0, "total_chunks": 2, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:1676" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s100_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task_PA.imagerobotics100K<n<1M0 likes394 downloads4mo agoHugging Face09Yanbin99 /Depth-Normal-Videos-42K Depth and Normal Videos Dataset 42,498 videos with depth and surface normals. Usage from huggingface_hub import hf_hub_download video = hf_hub_download( repo_id="Yanbin99/Depth-Normal-Videos-42K", filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4", repo_type="dataset" ) tabulardepth-estimation10K<n<100K1 likes332 downloads10mo agoHugging Face10N03N9 /cv24-tr-128-normalizedtext100K<n<1M0 likes305 downloads10mo agoHugging Face11winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes290 downloads1y agoHugging Face12kgnlp /meld-open-normalized MELD Open (Normalized) MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details. Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.tabulartoken-classification10M<n<100M0 likes251 downloads5mo agoHugging Face13csoai /gspc-normalized GSPC normalised — every bank in one schema The one schema to read first. 518 rows that flatten several GSPC banks into a single shape: source (the bank repository the row came from), axis, category, anchor, prompt, expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently dropped. If you want to reuse the banks without learning each one's native layout, start here. The live board is the authority GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.textothern<1K0 likes249 downloads5d agoHugging Face14mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes244 downloads11mo agoHugging Face15Scicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes233 downloads6mo agoHugging Face16ducido /merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 800, "total_frames": 133851, "total_tasks": 40, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:800" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task.imagerobotics100K<n<1M0 likes232 downloads7mo agoHugging Face17warmestman /common-voice-20-mn-normalized Common Voice 20.0 Mongolian Dataset This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release. Dataset Structure The dataset contains: Audio clips in .mp3 format Transcriptions for each audio clip Train/test/dev splits Additional metadata including speaker demographics Usage This dataset can be used for: Speech Recognition Voice Analysis Linguistic Research Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.audioautomatic-speech-recognition10K<n<100K4 likes181 downloads2y agoHugging Face18llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes180 downloads2y agoHugging Face19Twelve2five /igbo_tts_normalizedaudio100K<n<1M2 likes180 downloads1y agoHugging Face20syp1229 /E_normal_over70_add Dataset Card for "E_normal_over70_add" More Information needed tabular1K<n<10K0 likes178 downloads3y agoHugging Face21N03N9 /cv24-ru-128-normalizedtext100K<n<1M0 likes176 downloads10mo agoHugging Face22N03N9 /cv24-ur-128-normalizedtext10K<n<100K0 likes167 downloads10mo agoHugging Face23ChenWu98 /stack-v2-python-normal-onlytext1M<n<10M0 likes165 downloads1y agoHugging Face24N03N9 /cv24-cy-128-normalizedtext10K<n<100K0 likes161 downloads10mo agoHugging Face25izumi-lab /cc100-ja-filter-ja-normal Dataset Card for "cc100-ja-debug-filter-ja-normal" More Information needed text100M<n<1B2 likes158 downloads3y agoHugging Face26Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes156 downloads1mo agoHugging Face27N03N9 /cv24-de-128-normalizedtext100K<n<1M0 likes149 downloads10mo agoHugging Face28N03N9 /cv24-uk-128-normalizedtext10K<n<100K0 likes148 downloads10mo agoHugging Face29TechWolf /Skill-normalisation-ESCO-graded skill-normalisation-esco-graded Graded-relevance annotations for surface skill terms (ESCO alt-labels) from ESCO v1.1.0 skill-normalisation pairs against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 50 _id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.text1M<n<10M0 likes146 downloads2mo agoHugging Face30Kudod /VFD_normalize_9_v1tabular100K<n<1M0 likes141 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.