Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes733 downloads1y agoHugging Face02TongheZhangTH /XDof-TshirtFolding-20hours-normalizedtabular1M<n<10M0 likes592 downloads9mo agoHugging Face03winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes290 downloads1y agoHugging Face04kgnlp /meld-open-normalized MELD Open (Normalized) MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details. Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.tabulartoken-classification10M<n<100M0 likes251 downloads5mo agoHugging Face05mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes244 downloads11mo agoHugging Face06TongheZhangTH /CartonPickNPlace2Target-normalizedtabular10K<n<100K0 likes188 downloads9mo agoHugging Face07vc940 /business-entity-resolution-normalized Business Entity Resolution: normalised records Normalised copies of the six source files of the ML Challenge 2026 Business Entity Resolution task (business records from three sources, US / India in train, plus France in test). The goal of the task is to find, for every Source 1 record, the Source 2 / Source 3 records that describe the same business. File Rows train_s1.parquet 2,206,821 train_s2.parquet 5,034,616 train_s3.parquet 5,285,603 test_s1.parquet 1,732… See the full description on the dataset page: https://huggingface.co/datasets/vc940/business-entity-resolution-normalized.tabular10M<n<100M1 likes118 downloads16d agoHugging Face08Yiheyihe /galaxea-r1-shelf-full-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 99, "total_frames": 48085, "total_tasks": 1, "total_videos": 297, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:99" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Yiheyihe/galaxea-r1-shelf-full-normalized.tabularrobotics10K<n<100K0 likes81 downloads2y agoHugging Face09CarolinePascal /test_community_normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/test_community_normalized.tabularrobotics10K<n<100K0 likes79 downloads3mo agoHugging Face10winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk256-normalizedtabular10K<n<100K0 likes72 downloads1y agoHugging Face11CarolinePascal /test_community_normalized_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/test_community_normalized_1.tabularrobotics10K<n<100K0 likes70 downloads3mo agoHugging Face12jajostrains /Mathlib-Normalized-Sexpr Mathlib Normalized S-Expressions Lean 4 proof states from Mathlib, paired with the tactic applied at each step, in three representations extracted directly from the Lean kernel: Source-faithful S-expressions of the goal and every hypothesis, as Lean elaborated them. Normalized S-expressions of the same state, with stable local-context indices suitable for model input. Annotated tactic syntax -- the original tactic's syntax tree with identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.tabulartext-generation100K<n<1M0 likes65 downloads1mo agoHugging Face13Yiheyihe /galaxea-r1-shelf-10ep-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 11, "total_frames": 5218, "total_tasks": 1, "total_videos": 33, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:11" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Yiheyihe/galaxea-r1-shelf-10ep-normalized.tabularrobotics1K<n<10K0 likes61 downloads2y agoHugging Face14chcaa /dagw-word-frequencies-normalized-by-domain Dataset Card for DAGW Word Frequencies (normalized) Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421). Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com ) This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-normalized-by-domain.tabular10M<n<100M0 likes58 downloads4y agoHugging Face15Yiheyihe /galaxea-r1-shelf-1ep-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 1, "total_frames": 508, "total_tasks": 1, "total_videos": 3, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Yiheyihe/galaxea-r1-shelf-1ep-normalized.tabularroboticsn<1K0 likes51 downloads2y agoHugging Face16oms524 /place_spam_into_the_white_box_30hz_normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.images.front": { "dtype": "video", "shape": [ 480, 640, 3 ], "names": [ "height", "width", "channels" ], "info": { "is_depth_map":… See the full description on the dataset page: https://huggingface.co/datasets/oms524/place_spam_into_the_white_box_30hz_normalized.tabularrobotics10K<n<100K0 likes46 downloads2mo agoHugging Face17Yiheyihe /galaxea-r1-shelf-debug-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 1, "total_frames": 454, "total_tasks": 1, "total_videos": 3, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Yiheyihe/galaxea-r1-shelf-debug-normalized.tabularroboticsn<1K0 likes44 downloads2y agoHugging Face18CarolinePascal /test-community-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/test-community-normalized.tabularrobotics10K<n<100K0 likes42 downloads3mo agoHugging Face19martonbodo /rq4-pap-four-objects-275-hsv-sam-half-seed42-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 275, "total_frames": 92173, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:275" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martonbodo/rq4-pap-four-objects-275-hsv-sam-half-seed42-normalized.tabularrobotics10K<n<100K0 likes40 downloads4mo agoHugging Face20averrous /alljoined-normalized-subjecttabular10K<n<100K0 likes39 downloads1y agoHugging Face21martonbodo /pap-final-four-objects-550-cutie-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 550, "total_frames": 183669, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:550" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martonbodo/pap-final-four-objects-550-cutie-normalized.tabularrobotics100K<n<1M0 likes36 downloads4mo agoHugging Face22saipuneethgottam /sweep_dataset_merged_normalized_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 200, "total_frames": 137050, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:200" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/saipuneethgottam/sweep_dataset_merged_normalized_v2.tabularrobotics100K<n<1M0 likes34 downloads5mo agoHugging Face23saipuneethgottam /sweep_dataset_merged_normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 200, "total_frames": 137050, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:200" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/saipuneethgottam/sweep_dataset_merged_normalized.tabularrobotics100K<n<1M0 likes33 downloads5mo agoHugging Face24nlp-pw /Disaster-Tweets-Normalizedtabular100K<n<1M1 likes32 downloads3y agoHugging Face25benmayeux /hilserl_normalized_ee_deltaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 20, "total_frames": 6824, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:20" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/benmayeux/hilserl_normalized_ee_delta.tabularrobotics1K<n<10K0 likes32 downloads8mo agoHugging Face26benmayeux /hilserl_normalized_ee_delta_initialThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 5, "total_frames": 1041, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/benmayeux/hilserl_normalized_ee_delta_initial.tabularrobotics1K<n<10K0 likes32 downloads8mo agoHugging Face27justicedao /netherlands-laws-nl-normalized Netherlands Laws (Dutch, Normalized) Hugging Face target: justicedao/netherlands-laws-nl-normalized. This package is a normalized version of the Netherlands laws scrape output. This is a capped Netherlands scrape, not the full Dutch corpus. The scrape used max_documents=100, parsed 151 law record(s), and discovered 626 unique official BWBR law document(s) before applying the cap. Documents failed: 0. This refresh includes parser coverage improvements for older/French heading… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/netherlands-laws-nl-normalized.tabulartext-retrieval1K<n<10K0 likes32 downloads4mo agoHugging Face28ywlin /r1-pick-cup-stand-5x10eps-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 10, "total_frames": 12952, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ywlin/r1-pick-cup-stand-5x10eps-normalized.tabularrobotics10K<n<100K0 likes30 downloads1y agoHugging Face29saipuneethgottam /sweep_dataset_merged_normalized_clean80_from_sourcesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 80, "total_frames": 54820, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:80" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/saipuneethgottam/sweep_dataset_merged_normalized_clean80_from_sources.tabularrobotics10K<n<100K0 likes30 downloads5mo agoHugging Face30martonbodo /rq2-pap-red-block-100-hsv-sam-normalizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 100, "total_frames": 32126, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martonbodo/rq2-pap-red-block-100-hsv-sam-normalized.tabularrobotics10K<n<100K0 likes29 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.