datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drlc-leaderboard-dataCADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.criteo-attribution-dataset
Criteo Attribution Modeling for Bidding Dataset
This dataset is released along with the paper:
Attribution Modeling Increases Efficiency of Bidding in Display Advertising
Eustache Diemert*, Julien Meynet* (Criteo Research), Damien Lefortier (Facebook), Pierre Galland (Criteo) *authors contributed equally
2017 AdKDD & TargetAd Workshop, in conjunction with The 23rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2017)
When using this dataset, please cite the paper… See the full description on the dataset page: https://huggingface.co/datasets/criteo/criteo-attribution-dataset.data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/lukebarousse/data_jobs.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/4141ms/CADS-dataset.ai-model-popularity
Datamata AI Model Popularity Index
Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot.
Latest snapshot: 2026-10-04
Models in this release: 50
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology: https://www.datamatastudios.com/datasets
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT… See the full description on the dataset page: https://huggingface.co/datasets/mrmrx/CADS-dataset.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/WY-0206/CADS-dataset.hnm-fashion-recommendations-data
Dataset Rekomendasi Fashion H&M
Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam.
Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with 1440 full contest transcripts
of an LLM attempting them under simulated contest rules across three arms: with no hint,
with the official editorial as a hint, and with a hint written by a second model that gets
10 rounds of measured feedback to improve it.
Selection
The agent
Every contest run in this dataset comes from:… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/sunghong/CADS-dataset.predictive-stock-datasetOpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.angcb-data
Usage
python pipeline.py --window 6 --patch 18 --model cnn --disable_tqdm
window: the length of lookback months
patch: the size of small patch, should be divided by 180
model: cnn, rnn, cnn_rnn, currently cnn also captures the temporal correlation in a simple way, and demonstrates the most robust performance
--disable_tqdm: whether to show the progress bar
Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.openadmet-expansionrx-challenge-data
OpenADMET-ExpansionRx Challenge FULL dataset
This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026.
Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases.
While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.fsrs-datasetahr999-dataset
AHR999 BTC Hoarding Index Dataset
Open, daily-updated AHR999 BTC hoarding index dataset, self-computed from
Binance BTCUSDT daily closes and published as CSV and JSON.
This Hugging Face repository is a mirror. The canonical dataset endpoints are:
Dashboard: https://ahr999.aix4u.com/
GitHub: https://github.com/RuochenLyu/ahr999-dataset
CSV endpoint: https://ahr999.aix4u.com/datasets/ahr999.csv
JSON endpoint: https://ahr999.aix4u.com/datasets/ahr999.json
Kaggle discovery mirror:… See the full description on the dataset page: https://huggingface.co/datasets/kshift/ahr999-dataset.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.spartina-ai-eco-evolution-data
Spartina AI eco-evolutionary reanalysis
Code repository: github.com/ydchen0806/spartina-ai-eco-evolution
Reproducible code, derived tables and audit figures for the Spartina alterniflora aerial-observation project. The release reconstructs the available 2014–2021 patch-record analysis and audits the recovered 2014–2020 environment–growth simulation archive.
Active manuscript and validation pilot
The full working manuscript is available as
editable DOCX and
PDF. It… See the full description on the dataset page: https://huggingface.co/datasets/cyd0806/spartina-ai-eco-evolution-data.cve-and-cwe-dataset-1999-2025This collection brings together every Common Vulnerabilities & Exposures (CVE) entry published in the National Vulnerability Database (NVD) from the very first identifier — CVE-1999-0001 — through all records available on 30 May 2025.
It was built automatically with a Python script that calls the NVD REST API v2.0 page-by-page, handles rate-limits, and filters data.
After download each CVE object is pared down to the essentials and written to CVE_CWE_2025.csv with the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/stasvinokur/cve-and-cwe-dataset-1999-2025.CostNav-Teleop-Dataset
CostNav Teleop Dataset
Dataset Summary
The CostNav Teleop Dataset is a large-scale collection of human teleoperation recordings for robot navigation in an urban sidewalk simulation environment. It was collected as part of the CostNav benchmark, which evaluates navigation systems using real-world economic cost and revenue metrics rather than purely technical metrics.
The dataset contains 2,203 teleoperation episodes totaling 50.2 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/CostNav-Teleop-Dataset.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.PRCA-Net-dataset
Ray-Traced Cross-Frequency Radio Map Dataset
A large ray-traced radio-map (path-loss) dataset for zero-shot
cross-frequency generalization research, generated with
Sionna RT over real urban geometry from
OpenStreetMap.
150 urban scenes across 15 cities, 256×256 rasters
8 transmitters per scene across three deployment strata (street,
rooftop, mast)
6 carrier frequencies: 1.8, 3.5, 7, 28 GHz (training) + 10, 60 GHz
(held out, for interpolation / extrapolation studies)
7,200… See the full description on the dataset page: https://huggingface.co/datasets/SHussain37/PRCA-Net-dataset.scriptscryptocurrency-futures-ohlcv-dataset-1m
