Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /DatasetWithCapitalLetterstextn<1K0 likes37k downloads3y agoHugging Face02huggingface-projects /drlc-leaderboard-datatabular10K<n<100K2 likes24k downloads8d agoHugging Face03datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes19k downloads3y agoHugging Face04huggingface /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.tabularimage-segmentation10K<n<100K4 likes17k downloads10mo agoHugging Face05SaProtHub /Dataset-GB1-fitness Description This dataset contains fitness score of mutant GB1 protein. Protein Format: AA sequence Splits traing: 119644 valid: 14917 test: 14800 Related paper Nicholas C Wu, Lei Dai, C Anders Olson, James O Lloyd-Smith, Ren Sun (2016) Adaptation in protein fitness landscapes is facilitated by indirect paths eLife 5:e16965 https://doi.org/10.7554/eLife.16965 Label Label is the fitness of mutant protein. The fitness of each variant can be viewed as… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-GB1-fitness.text100K<n<1M2 likes13k downloads2y agoHugging Face06bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K199 likes11k downloads2y agoHugging Face07databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K30 likes9k downloads3mo agoHugging Face08TACK-project /TACK_Tunnel_Data TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/TACK-project/TACK_Tunnel_Data.image1K<n<10K12 likes8.5k downloads10mo agoHugging Face09lukebarousse /data_jobs 🧠 data_jobs Dataset A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse. Background I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources. You can find the full dataset at my app datanerd.tech. Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/lukebarousse/data_jobs.tabular100K<n<1M101 likes7.7k downloads1y agoHugging Face104141ms /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/4141ms/CADS-dataset.tabularimage-segmentation10K<n<100K0 likes7.1k downloads10mo agoHugging Face11Abtinzandi /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K18 likes6.4k downloads5d agoHugging Face12datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-10-04 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes5.8k downloads6d agoHugging Face13maharshipandya /spotify-tracks-dataset Content This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly. Usage The dataset can be used for: Building a Recommendation System based on some user input or preference Classification purposes based on audio features and available genres Any other application that you can think of. Feel free to discuss! Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.tabularfeature-extraction100K<n<1M136 likes4.7k downloads3y agoHugging Face14MVU-Eval-Team /MVU-Eval-Data MVU-Eval Dataset Paper | Code | Project Page Dataset Description The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.tabularvideo-text-to-text1K<n<10K2 likes3.8k downloads11mo agoHugging Face15arekborucki /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.tabularimage-segmentation10K<n<100K2 likes3.4k downloads10mo agoHugging Face16datasets-examples /doc-formats-csv-1 [doc] formats - csv - 1 This dataset contains one csv file at the root: data.csv kind,sound dog,woof cat,meow pokemon,pika human,hello The YAML section of the README does not contain anything related to loading the data (only the size category metadata): --- size_categories: - n<1K --- textn<1K0 likes3.2k downloads3y agoHugging Face17mrmrx /CADS-datasetgated CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT… See the full description on the dataset page: https://huggingface.co/datasets/mrmrx/CADS-dataset.tabularimage-segmentation10K<n<100K58 likes3.2k downloads4mo agoHugging Face18WY-0206 /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/WY-0206/CADS-dataset.tabularimage-segmentation10K<n<100K0 likes3.1k downloads9mo agoHugging Face19jingwei-xu-00 /eccv2026-cad-challenge-data ECCV 2026 CAD Challenge Data This challenge is part of the workshop The Path to Manufacturing: Evolving 3D Generation to Intelligent Computer-Aided Design. Workshop homepage: https://3dgen-cad-workshop.github.io/ Challenge submission Space: https://huggingface.co/spaces/jingwei-xu-00/eccv2026-cad-challenge Dataset rendering and preparation code (only .step files are required): https://github.com/DavidXu-JJ/eccv2026-cad-challenge-data-render This repository contains the public… See the full description on the dataset page: https://huggingface.co/datasets/jingwei-xu-00/eccv2026-cad-challenge-data.3dimage-to-3d1K<n<10K7 likes3.1k downloads2mo agoHugging Face20einrafh /hnm-fashion-recommendations-data Dataset Rekomendasi Fashion H&M Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam. Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.imagetabular-classification10M<n<100M4 likes2.9k downloads1y agoHugging Face21INV-WZQ /ReactiveGWM-Datasets ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models 📚 Datasets-Introduction ReactiveGWM-Datasets is the strategy-aligned training corpus that powers ReactiveGWM, a game world model that decouples player control from NPC autonomy. To learn that decoupling, the model needs supervision that pairs each gameplay clip with both a per-frame action stream (what the player did) and a high-level NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.textimage-to-video10K<n<100K8 likes2.9k downloads5mo agoHugging Face22bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes2.8k downloads2y agoHugging Face23xupy21 /ICPC_Data ICPC World Finals — a discriminative subset, with model traces 24 ICPC World Finals problems (2021–2025), together with 1440 full contest transcripts of an LLM attempting them under simulated contest rules across three arms: with no hint, with the official editorial as a hint, and with a hint written by a second model that gets 10 rounds of measured feedback to improve it. Selection The agent Every contest run in this dataset comes from:… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.tabulartext-generationn<1K1 likes2.5k downloads9d agoHugging Face24sunghong /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/sunghong/CADS-dataset.tabularimage-segmentation10K<n<100K0 likes2.4k downloads10mo agoHugging Face25bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes2.2k downloads2y agoHugging Face26nadtoka /predictive-stock-datasettabular1K<n<10K0 likes2.2k downloads18h agoHugging Face27heispv /protein_data_testsplit 1, 2 -> for sequences split 3, 4 -> for residues textn<1K0 likes2.2k downloads2y agoHugging Face28zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes2.1k downloads3y agoHugging Face29ShiroOnigami23 /skin-cancer-ham10000-datasetimage1K<n<10K1 likes1.8k downloads9mo agoHugging Face30RepDB /exercise-dataset Exercise Dataset — Free Tier (RepDB) A free, ready-to-use fitness exercise dataset: 609 exercises, each illustrated with flat-style 512×512 WebP images (a start/peak pose pair, or a single main pose for static holds and stretches), with target muscles, equipment, MET values, and full instructions in English, German, and Spanish. This public snapshot is the free tier of RepDB. Free for personal and commercial use inside applications, with attribution. Need exercise… See the full description on the dataset page: https://huggingface.co/datasets/RepDB/exercise-dataset.imagen<1K4 likes1.8k downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.