datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.origindata-preview-sample
OriginData (preview sample)
OriginData is the world's first large-scale, real-world dataset combining hand pose and force annotations. Spanning 28 domains and 1,170 real-world tasks, it captures how human hands interact with the physical world, providing force, pose, and semantic annotations supported by high-precision multimodal calibration for embodied AI.
This preview contains 100.62 hours across 7,717 episodes, delivered in LeRobot v3.0 format with stereo RGB video, hand… See the full description on the dataset page: https://huggingface.co/datasets/OriginFlow-AI/origindata-preview-sample.fine-news-sample
Fine-News Sample
Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus.
The sample covers all 117 capture months and 388 language-and-script labels in that corpus.
Each selected row preserves its article text, source metadata, and sampling weight.
The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives.
At a glance
Measure
Value
Rows
1,000,000
Distinct document IDs
1,000,000
Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.Amazon_Sample_Metadata_2023
Dataset Card for Dataset Name
Original datasets can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset was made as sample of several datasets from the link above.
Dataset Description
This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors,
Health and Personal Care, Amazon Clothing Shoes and Jewlery,
Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.dolma3_300B_sample
Dolma 3 — 300B-token sample
🌐 The Fin AI
Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai.
Source
A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0.
Structure
Rows: 187,823,645
Columns: source, date, text, token_count, category
Quick Start
from datasets import load_dataset
ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.dolma_18bn_stratified_sampletweets_sample_2026
Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus
A large, deliberately untargeted sample of public posts from X/Twitter, collected via
Nitter by sweeping a 65,689-term newspaper vocabulary
rather than a topical keyword set.
It is built as a background / reference corpus: a baseline of "what was being said in
general" against which a topically targeted collection can be contrasted. It is the reference
arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.jade-samples-10000x1
JADE amortized posterior samples — 10,000 observations x 1 draw
Noisy weak-lensing convergence observations paired with joint posterior draws of
(convergence field, cosmology) from the amortized conditional diffusion model
of JADE.
[!IMPORTANT]
This dataset is not a product of arXiv:2606.31988.
It was generated afterwards, with the same trained model, to support posterior
calibration diagnostics that do not appear in the paper. No number in the paper
was computed from it, and… See the full description on the dataset page: https://huggingface.co/datasets/b-remy/jade-samples-10000x1.dclm-refinedweb-600m-sampleegocentric-kitchen-sample
Diffraction Egocentric Kitchen Capture Sample
A small, inspectable sample of human kitchen manipulation captured with Stray Scanner on a LiDAR-equipped iPhone: native RGB, metric depth and confidence, per-frame camera calibration, device odometry, raw device IMU, and explicitly estimated hand/object annotations.
Human observation sample. License: cc-by-4.0. This sample contains 3 recordings totaling 167.85 seconds. It is an observation dataset for evaluating human-video… See the full description on the dataset page: https://huggingface.co/datasets/diffracting/egocentric-kitchen-sample.historical-futures-data-sample
Historical Futures Data Sample
This repository contains a free evaluation sample of historical futures data across selected contracts and frequencies.
The complete catalog covers more than 2,000 futures roots and 900 million observations.
View Data and Pricing: https://futuresforexandsomeindexes.com/
This package is a normalized evaluation sample containing 40 selected contracts across 8 futures roots: CL, ES, GC, SB, SR3, VX, ZC, ZN.
The original root and contract files are… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-futures-data-sample.fineweb-edu-sample-10BT-tiktokenizedsample_recovery-demonstrationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_widowxai_follower_robot",
"total_episodes": 60,
"total_frames": 53886,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/sample_recovery-demonstration.tokenized_samplehistorical-equity-index-futures-data-sample
Historical Equity Index Futures Data Sample
A free evaluation sample of historical equity index futures data covering selected U.S. and international benchmark index contracts.
Full historical futures catalog, broader contract coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/
This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is not… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-equity-index-futures-data-sample.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment.
kwaiklear-sample-level-agent-trajectories-2.2Mdental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Full-Mouth Three-Camera Sample)
A public preview of one complete full-mouth implant rehabilitation — both jaws in a
single session, six implants in the maxilla and six in the mandible with multi-unit
abutments — recorded by three synchronised cameras in a working operating room, with
the surgeons' own words aligned to the picture and every keyframe annotated across six
layers (L0–L5).
This repository is a showcase slice of… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.fineweb2_ar_65m_sampleclinical_trials_history_sample
Clinical Trials Version History: Sample
A free sample of the Clinical Trials Version History dataset by
Anamnesis Data: structured medical and scientific data for biotech,
pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or
when a sponsor moved a completion date.
The sample holds 20 complete trials (741 versions), one Parquet file per table
(79 tables… See the full description on the dataset page: https://huggingface.co/datasets/anamnesis-data/clinical_trials_history_sample.fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5
FineWeb-edu 10BT Sample embedded with nomic-text-v1.5
The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT.
The chunks were then embedded using nomic-text-v1.5.
Dataset Details
Dataset Sources
Repository: https://github.com/enjalot/fineweb-modal
Uses
Direct Use
The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs.
Files:
HRBench4K.tsv — 300 rows
HRBench8K.tsv — 300 rows
MathVision_MINI.tsv — 300 rows
MathVista_MINI.tsv — 300 rows
MMBench_en_dev.tsv — 300 rows
MME_RealWorld_Lite.tsv — 300 rows
MMMU_val.tsv — 300 rows
MMStar.tsv — 300 rows
MMVet.tsv — 218 rows
POPE.tsv — 300 rows
RealworldQA.tsv — 300 rows
SEED_Bench.tsv — 300 rows
VStarBench.tsv — 191 rows
Notes:
The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.dolma_18bn_prop_stratified_samplemanus-egocentric-sample
manus-egocentric-sample
Egocentric video dataset with Manus glove hand tracking data, converted to LeRobot v3.0 format.
Dataset Description
This dataset contains egocentric (first-person view) recordings of human hands performing various manipulation tasks, captured with:
Manus Metagloves: High-precision finger tracking (~70Hz)
OAK-D Camera: RGB video (1920x1080, 30fps) + Depth (640x400, 30fps)
IMU: Accelerometer and gyroscope data
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/OpenGraphLabs-Research/manus-egocentric-sample.korea-household-egocentric-samples
60BASE · Korea Household Egocentric Samples
Twenty household video excerpts from 60BASE's Korea household collection, including four additions prepared on September 27, 2026.
Use this sample pack to inspect the footage and discuss a full-recording request or a custom collection brief.
Discuss a data project · 60BASE · Email
Included in this release
Specification
Video
20 MP4 clips × 20 seconds; 6 minutes 40 seconds total
Tasks
Dishwashing, clothes/towel folding… See the full description on the dataset page: https://huggingface.co/datasets/60base/korea-household-egocentric-samples.
