datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fishnet-lichess-normalizednormalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.T2I-ImageNet-Normalsentence-level-detection-normal-v1Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans
Multimodal Chest X ray dataset for Normal and Bacterial Pneumonia in Africans | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: imagefolder - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans.mc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
modelnet40_normal_resampled-compressedOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedperturbseq_normalized
CellClip public normalized perturbation single-cell release
This public repository contains the source-cleared, human, cell-level portion
of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and
260,859,318,437 payload bytes. These are processed derivatives rather than
the original raw download archives.
Source partition
Units
Cells
Reference cells
Effect cells
scPerturb (23 genepert + 229 chempert)
252
4,774,802
397,528
4,377,274
XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.DiffusionPDE-normalizedWe take the dataset from DiffusionPDE. For convenience, we provide our processed version on Hugging Face (see Appendix D & E in our FunDPS paper for details). The processing scripts are provided under utils/ here. It is worth noting that we normalized the datasets to zero mean and 0.5 standard deviation to follow EDM's practice.
XDof-TshirtFolding-20hours-normalizedoscar2301-ja-filter-ja-normal
Dataset Card for "oscar2301-ja-filter-ja-normal"
More Information needed
merged_libero_s100_mdnl_original_s100_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1676,
"total_frames": 269918,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1676"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s100_mdnl_original_s100_synthetic_20ep_2_hard_task_5ep_all_normal_task.merged_libero_s100_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task_PAThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1676,
"total_frames": 269918,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1676"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s100_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task_PA.Depth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
cv24-tr-128-normalizedOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizednormal_test_benchskin-disease-acne-rosacea-normalmeld-open-normalized
MELD Open (Normalized)
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.gspc-normalized
GSPC normalised — every bank in one schema
The one schema to read first. 518 rows that flatten several GSPC banks into a
single shape: source (the bank repository the row came from), axis, category, anchor, prompt,
expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently
dropped. If you want to reuse the banks without learning each one's native layout, start here.
The live board is the authority
GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.lichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.Normalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 800,
"total_frames": 133851,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task.normal_ball_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 44089,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/satvikahuja/normal_ball_full.CartonPickNPlace2Target-normalizedcommon-voice-20-mn-normalized
Common Voice 20.0 Mongolian Dataset
This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release.
Dataset Structure
The dataset contains:
Audio clips in .mp3 format
Transcriptions for each audio clip
Train/test/dev splits
Additional metadata including speaker demographics
Usage
This dataset can be used for:
Speech Recognition
Voice Analysis
Linguistic Research
Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized
상세
데이터셋 설명
OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다.
OpenAI gpt-4o-mini를 통해 번역됐습니다.
Shared by llami-team
Language(s) (NLP): Korean
Uses
한국어 reasoning 모델 distillation
reasoning cold-start 데이터셋
Dataset Structure
question: 질문
reasoning: 추론 과정
response: 응답
Dataset Creation
[LLAMI Team] (https://llami.net)
LLAMI Github
lemon-mint
Source Data
OpenThoughts-114k-Normalized
igbo_tts_normalizedepsilon-normalized
Dataset Card for "epsilon-normalized"
More Information needed
