Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face02OriginFlow-AI /origindata-preview-samplegated OriginData (preview sample) OriginData is the world's first large-scale, real-world dataset combining hand pose and force annotations. Spanning 28 domains and 1,170 real-world tasks, it captures how human hands interact with the physical world, providing force, pose, and semantic annotations supported by high-precision multimodal calibration for embodied AI. This preview contains 100.62 hours across 7,717 episodes, delivered in LeRobot v3.0 format with stereo RGB video, hand… See the full description on the dataset page: https://huggingface.co/datasets/OriginFlow-AI/origindata-preview-sample.tabularrobotics10M<n<100M15 likes15k downloads1d agoHugging Face03ksolovev /fine-news-sample Fine-News Sample Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus. The sample covers all 117 capture months and 388 language-and-script labels in that corpus. Each selected row preserves its article text, source metadata, and sampling weight. The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives. At a glance Measure Value Rows 1,000,000 Distinct document IDs 1,000,000 Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.texttext-generation1M<n<10M0 likes2.7k downloads2d agoHugging Face04stablellama /Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source The images were created in ComfyUI with the bf16 version of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.tabulartext-to-image1K<n<10K0 likes2k downloads1mo agoHugging Face05ExylosAi /egocentric-vr-capture-20h-multimodal-sample Egocentric VR Capture — 20-Hour Multimodal Inspection Sample 195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.tabularrobotics1M<n<10M0 likes1.6k downloads18d agoHugging Face06smartcat /Amazon_Sample_Metadata_2023 Dataset Card for Dataset Name Original datasets can be found on: https://amazon-reviews-2023.github.io/ Dataset Details This dataset was made as sample of several datasets from the link above. Dataset Description This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors, Health and Personal Care, Amazon Clothing Shoes and Jewlery, Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.tabular1M<n<10M1 likes1.2k downloads2y agoHugging Face07TheFinAI /dolma3_300B_samplegated Dolma 3 — 300B-token sample 🌐 The Fin AI Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. Source A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. Structure Rows: 187,823,645 Columns: source, date, text, token_count, category Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.tabulartext-generation100M<n<1B0 likes1.1k downloads3d agoHugging Face08SBMM75 /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes961 downloads28d agoHugging Face09orionweller /dolma_18bn_stratified_sampletabular10M<n<100M0 likes795 downloads2y agoHugging Face10SinclairSchneider /tweets_sample_2026 Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus A large, deliberately untargeted sample of public posts from X/Twitter, collected via Nitter by sweeping a 65,689-term newspaper vocabulary rather than a topical keyword set. It is built as a background / reference corpus: a baseline of "what was being said in general" against which a topically targeted collection can be contrasted. It is the reference arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.tabulartext-classification10M<n<100M0 likes775 downloads2mo agoHugging Face11b-remy /jade-samples-10000x1 JADE amortized posterior samples — 10,000 observations x 1 draw Noisy weak-lensing convergence observations paired with joint posterior draws of (convergence field, cosmology) from the amortized conditional diffusion model of JADE. [!IMPORTANT] This dataset is not a product of arXiv:2606.31988. It was generated afterwards, with the same trained model, to support posterior calibration diagnostics that do not appear in the paper. No number in the paper was computed from it, and… See the full description on the dataset page: https://huggingface.co/datasets/b-remy/jade-samples-10000x1.tabular10K<n<100K0 likes773 downloads1mo agoHugging Face12blab-jhu /dclm-refinedweb-600m-sampletabular100M<n<1B0 likes771 downloads2mo agoHugging Face13diffracting /egocentric-kitchen-sample Diffraction Egocentric Kitchen Capture Sample A small, inspectable sample of human kitchen manipulation captured with Stray Scanner on a LiDAR-equipped iPhone: native RGB, metric depth and confidence, per-frame camera calibration, device odometry, raw device IMU, and explicitly estimated hand/object annotations. Human observation sample. License: cc-by-4.0. This sample contains 3 recordings totaling 167.85 seconds. It is an observation dataset for evaluating human-video… See the full description on the dataset page: https://huggingface.co/datasets/diffracting/egocentric-kitchen-sample.imagen<1K0 likes691 downloads1mo agoHugging Face14lynx1231 /historical-futures-data-sample Historical Futures Data Sample This repository contains a free evaluation sample of historical futures data across selected contracts and frequencies. The complete catalog covers more than 2,000 futures roots and 900 million observations. View Data and Pricing: https://futuresforexandsomeindexes.com/ This package is a normalized evaluation sample containing 40 selected contracts across 8 futures roots: CL, ES, GC, SB, SR3, VX, ZC, ZN. The original root and contract files are… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-futures-data-sample.tabular1M<n<10M1 likes597 downloads2mo agoHugging Face15snimu /fineweb-edu-sample-10BT-tiktokenizedtabular1M<n<10M0 likes573 downloads2y agoHugging Face16REBOOT26 /sample_recovery-demonstrationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_widowxai_follower_robot", "total_episodes": 60, "total_frames": 53886, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:60" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/sample_recovery-demonstration.tabularrobotics10K<n<100K0 likes544 downloads5mo agoHugging Face17Pclanglais /tokenized_sampletabular1M<n<10M0 likes543 downloads2y agoHugging Face18lynx1231 /historical-equity-index-futures-data-sample Historical Equity Index Futures Data Sample A free evaluation sample of historical equity index futures data covering selected U.S. and international benchmark index contracts. Full historical futures catalog, broader contract coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/ This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is not… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-equity-index-futures-data-sample.tabular100K<n<1M0 likes479 downloads22d agoHugging Face19stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes471 downloads1mo agoHugging Face20BAAI /RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment. tabularn<1K2 likes453 downloads1y agoHugging Face21israel-adewuyi /kwaiklear-sample-level-agent-trajectories-2.2Mtabular1M<n<10M0 likes448 downloads7mo agoHugging Face22OralSurgery /dental-implant-surgery-sample Dental Implant Surgery — Multimodal Annotated Video (Full-Mouth Three-Camera Sample) A public preview of one complete full-mouth implant rehabilitation — both jaws in a single session, six implants in the maxilla and six in the mandible with multi-unit abutments — recorded by three synchronised cameras in a working operating room, with the surgeons' own words aligned to the picture and every keyframe annotated across six layers (L0–L5). This repository is a showcase slice of… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.imagevideo-classification1K<n<10K1 likes444 downloads4d agoHugging Face23akhooli /fineweb2_ar_65m_sampletabular1M<n<10M0 likes417 downloads2y agoHugging Face24anamnesis-data /clinical_trials_history_sample Clinical Trials Version History: Sample A free sample of the Clinical Trials Version History dataset by Anamnesis Data: structured medical and scientific data for biotech, pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or when a sponsor moved a completion date. The sample holds 20 complete trials (741 versions), one Parquet file per table (79 tables… See the full description on the dataset page: https://huggingface.co/datasets/anamnesis-data/clinical_trials_history_sample.tabulartabular-classification100K<n<1M2 likes417 downloads11d agoHugging Face25enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes408 downloads2y agoHugging Face26Icey444 /tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs. Files: HRBench4K.tsv — 300 rows HRBench8K.tsv — 300 rows MathVision_MINI.tsv — 300 rows MathVista_MINI.tsv — 300 rows MMBench_en_dev.tsv — 300 rows MME_RealWorld_Lite.tsv — 300 rows MMMU_val.tsv — 300 rows MMStar.tsv — 300 rows MMVet.tsv — 218 rows POPE.tsv — 300 rows RealworldQA.tsv — 300 rows SEED_Bench.tsv — 300 rows VStarBench.tsv — 191 rows Notes: The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.tabularn<1K0 likes398 downloads3mo agoHugging Face27voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face28orionweller /dolma_18bn_prop_stratified_sampletabular10M<n<100M0 likes370 downloads2y agoHugging Face29OpenGraphLabs-Research /manus-egocentric-sample manus-egocentric-sample Egocentric video dataset with Manus glove hand tracking data, converted to LeRobot v3.0 format. Dataset Description This dataset contains egocentric (first-person view) recordings of human hands performing various manipulation tasks, captured with: Manus Metagloves: High-precision finger tracking (~70Hz) OAK-D Camera: RGB video (1920x1080, 30fps) + Depth (640x400, 30fps) IMU: Accelerometer and gyroscope data Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/OpenGraphLabs-Research/manus-egocentric-sample.tabularrobotics10K<n<100K1 likes367 downloads9mo agoHugging Face3060base /korea-household-egocentric-samples 60BASE · Korea Household Egocentric Samples Twenty household video excerpts from 60BASE's Korea household collection, including four additions prepared on September 27, 2026. Use this sample pack to inspect the footage and discuss a full-recording request or a custom collection brief. Discuss a data project · 60BASE · Email Included in this release Specification Video 20 MP4 clips × 20 seconds; 6 minutes 40 seconds total Tasks Dishwashing, clothes/towel folding… See the full description on the dataset page: https://huggingface.co/datasets/60base/korea-household-egocentric-samples.tabularn<1K1 likes365 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.