Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01transferable-samplers /many-peptides-md [!IMPORTANT] Critical Update The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper. Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results. Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26. Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.12 likes1.8m downloads10mo agoHugging Face02ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K63 likes1.1m downloads1y agoHugging Face03malcolmrey /samplesimage1K<n<10K10 likes101k downloads4d agoHugging Face04P2SAMAPA /p2-etf-samba-models0 likes43k downloads6h agoHugging Face05m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face06blueyo0 /SA-Med3D-140K SA-Med3D-140K [github] Dataset Summary SA-Med3D-140K is a large-scale, multi-modal, multi-anatomical volumetric medical image segmentation dataset. It was created to facilitate the development of general-purpose foundation models for 3D medical image segmentation. The dataset comprises 21,729 3D medical images and 143,518 corresponding masks. It was gathered from a combination of 70 public datasets and 8,128 privately licensed annotated cases from 24 hospitals.… See the full description on the dataset page: https://huggingface.co/datasets/blueyo0/SA-Med3D-140K.11 likes19k downloads1y agoHugging Face07hf-internal-testing /dummy-audio-samplesaudion<1K0 likes16k downloads19d agoHugging Face08axolotl-ai-co /evolkit-logprobs-pipeline-75k-v2-sampletextn<1K1 likes15k downloads2y agoHugging Face09OriginFlow-AI /origindata-preview-samplegated OriginData (preview sample) OriginData is the world's first large-scale, real-world dataset combining hand pose and force annotations. Spanning 28 domains and 1,170 real-world tasks, it captures how human hands interact with the physical world, providing force, pose, and semantic annotations supported by high-precision multimodal calibration for embodied AI. This preview contains 100.62 hours across 7,717 episodes, delivered in LeRobot v3.0 format with stereo RGB video, hand… See the full description on the dataset page: https://huggingface.co/datasets/OriginFlow-AI/origindata-preview-sample.tabularrobotics10M<n<100M15 likes15k downloads1d agoHugging Face10zhangtao-whu /sam_tfrecords0 likes12k downloads1y agoHugging Face11ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes12k downloads2y agoHugging Face12Vchitect /VBench-2.0_sampled_videos Sample Videos of VBench-2.0 This dataset is used in the paper:👉 arXiv:2503.21755 video10K<n<100K0 likes11k downloads1y agoHugging Face13OpenGVLab /SA-Med2D-20M SA-Med2D-20M The largest benchmark dataset for segmentation in the field of medical imaging. As is well known, the emergence of ImageNet has greatly propelled the development of AI, especially deep learning. It has provided massive data and powerful baseline models for the computer vision community, enabling researchers to achieve breakthroughs in tasks such as natural image classification, segmentation, and detection. However, in the medical image realm, there lack of such a large… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/SA-Med2D-20M.27 likes11k downloads3y agoHugging Face14sayakpaul /sample-datasetsimagen<1K1 likes11k downloads3mo agoHugging Face15Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes11k downloads3mo agoHugging Face16TIACentre /TIAToolBox_Remote_Samples LICENSE No re-distribution allowed. Purpose This repository contains publicly available samples used by the TIAToolBox for testing purposes. Some of these images have been downloaded from [OpenSlide] for code verification purposes. GitHub Repository: [TIAToolBox] other0 likes10k downloads21h agoHugging Face17moonshine-ai /audio_samples_1kaudio0 likes9k downloads7mo agoHugging Face18zeahub /camus-sample CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus. This dataset is a zea-format (HDF5) conversion of the CAMUS dataset for multi-structure segmentation in 2-D echocardiography. Property Value Modality 2-D transthoracic echocardiography Patients 500 Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.image-segmentation0 likes8.8k downloads2mo agoHugging Face19Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8k downloads1y agoHugging Face20samuelt7 /csf0 likes7.8k downloads3mo agoHugging Face21icyCreater /SpreadsheetBench-2-sample SpreadsheetBench 2 Dataset SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization. Overview Directory Task type Tasks Input workbooks Gold workbooks Debugging Spreadsheet debugging and error correction 100 100 10 Financial_Model Completion of multi-sheet financial models 100 100 20 Template Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.0 likes7.2k downloads3mo agoHugging Face22sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes7.2k downloads9mo agoHugging Face23Sam04 /au30_tra Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Sam04/au30_tra.0 likes7k downloads8mo agoHugging Face24unileon-robotics /malware-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository contains real samples of malware that can be executed (.exe) and artifacts related with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). DO NOT execute any of… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/malware-samples.1K<n<10K7 likes6.6k downloads2mo agoHugging Face25Vchitect /VBench-I2V_sampled_videovideon<1K0 likes6.5k downloads7mo agoHugging Face26spmoisa548 /sama11 likes6.3k downloads5d agoHugging Face27samelias1 /Helium_TG0 likes6.2k downloads5d agoHugging Face28samuelt7 /lsd0 likes5.8k downloads1y agoHugging Face29paxini /Omnisharing_DB_SampleData Overview The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios. The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.8 likes5.8k downloads5mo agoHugging Face30knkarthick /samsum Dataset Card for SAMSum Corpus Dataset Description Links Homepage: hhttps://arxiv.org/abs/1911.12237v2 Repository: https://arxiv.org/abs/1911.12237v2 Paper: https://arxiv.org/abs/1911.12237v2 Point of Contact: https://huggingface.co/knkarthick Dataset Summary The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.textsummarization10K<n<100K44 likes5.5k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.