Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Robeedau /airlens-live AirLens Live Data Live data layer for AirLens, an open air-quality monitoring platform. Updated by scheduled GitHub Actions pipelines. Layout mirrors the former Supabase Storage buckets: Path Content Cadence aq-data/current-*-grid.json Global pollutant grids (PM2.5/PM10/O3/NO2/CO) hourly aq-data/timeline/ GEFS-Aerosols PM2.5 frames, -24h..+24h, 3h step every 3h aq-data/predictions/grid_latest.json AOD→PM2.5 model predictions (p10-p90 + DQSS) every 3h… See the full description on the dataset page: https://huggingface.co/datasets/Robeedau/airlens-live.tabular100K<n<1M2 likes63k downloads48m agoHugging Face02lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips 37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio. Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video. 🚀 Explore it in LightlyStudio hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips cd epic-kitchens-100-clips pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K3 likes17k downloads4d agoHugging Face03di-zhang-fdu /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. Citation @misc {di_zhang_2025, author = { {Di Zhang} }, title = { AIME_1983_2024 (Revision 6283828) }, year = 2025, url = {… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/AIME_1983_2024.tabularn<1K41 likes16k downloads2y agoHugging Face04Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K5 likes8.9k downloads5mo agoHugging Face05stanford-crfm /air-bench-2024 AIRBench 2024 AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning categories of the regulation-based safety categories in the AIR 2024 safety taxonomy. Dataset Details Dataset Description AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.texttext-generation10K<n<100K26 likes6k downloads2y agoHugging Face06datamatastudios /ai-model-popularity Datamata AI Model Popularity Index Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot. Latest snapshot: 2026-10-04 Models in this release: 50 Updated: weekly Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution. Source & methodology: https://www.datamatastudios.com/datasets Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.tabularn<1K0 likes5.8k downloads6d agoHugging Face07APRIL-AIGC /UltraVideo UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.tabularimage-to-video10K<n<100K68 likes4.6k downloads1y agoHugging Face08gneubig /aime-1983-2024 AIME Problem Set 1983-2024 Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) from 1983 to 2024. The AIME is a prestigious mathematics competition for high school students in the United States and Canada. Dataset Summary Source: Kaggle - AIME Problem Set 1983-2024 License: CC0: Public Domain Total Problems: 2,250 Years Covered: 1983 to 2024 Main Task: Mathematics Problem Solving… See the full description on the dataset page: https://huggingface.co/datasets/gneubig/aime-1983-2024.tabulartext-classificationn<1K21 likes4.5k downloads2y agoHugging Face09presentofai /ai-timeline AI Timeline Dataset An open, dated, source-linked record of what artificial intelligence actually did between July 2025 and today. Every row is a single real-world event with a primary source attached. 3,473 events · 934 distinct publishers · 2025-07-01 to 2026-10-10 Maintained by Present of AI, a daily AI news site. Updated as the timeline grows. Why this exists Most AI datasets are benchmarks or model outputs. This one is a record of events: deployments, funding… See the full description on the dataset page: https://huggingface.co/datasets/presentofai/ai-timeline.tabulartext-classification1K<n<10K0 likes4.2k downloads5h agoHugging Face10handshake-ai-research /studentbench StudentBench StudentBench: AI and human tutoring yield equivalent GRE learning gains Paper · Reproduction code · Project · Files StudentBench measures how well AI tutors help real students learn. We release the study data to reproduce the paper's results and support open research on learning, lesson planning, practice problems, tutoring conversations, engagement and cost. Pooled AI tutoring and expert human tutoring produced equivalent GRE learning gains ($p = .015$). Six AI… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/studentbench.tabularn<1K3 likes3.5k downloads12d agoHugging Face11Censius-AI /ECommerce-Women-Clothing-Reviewstabular10K<n<100K2 likes3.4k downloads4y agoHugging Face12lmarena-ai /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.tabulartext-classification10K<n<100K159 likes2.4k downloads2y agoHugging Face13AI4Sec /cti-bench Dataset Card for CTIBench A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks. Dataset Details Dataset Description CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI. Components: CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.textzero-shot-classification1K<n<10K21 likes2.3k downloads2y agoHugging Face14APRIL-AIGC /UltraVideo-Long UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.tabularimage-to-video10K<n<100K7 likes2.3k downloads1y agoHugging Face15Ahus-AIM /EchoXFlow EchoXFlow This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam. Extraction Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder: git lfs install git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow cd EchoXFlow The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam archive into a local data/ directory to materialize the… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/EchoXFlow.textn<1K5 likes2k downloads5mo agoHugging Face16vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K11 likes1.9k downloads1y agoHugging Face17Beijing-AISI /panda-bench PandaBench PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies. The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges. Dataset Description This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.tabulartext-generation100K<n<1M0 likes1.9k downloads1y agoHugging Face18AImageLab-Zip /mimose_runstabular1K<n<10K0 likes1.8k downloads3d agoHugging Face19aicostbudget-ai /ai-api-pricing AI API Pricing Dataset Source-linked AI API pricing data covering token, cache, batch, tiered, multimodal, and non-token pricing across multiple providers, including OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and Cohere. These are examples, not an exhaustive provider list. Live dataset and documentation Source repository Fixed v1.1.0 release (snapshot 2026-09-30) Version DOI Methodology This Hugging Face dataset is the machine-readable distribution of the public AI API… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.tabularn<1K0 likes1.8k downloads4h agoHugging Face20genbio-ai /rna-downstream-tasks GB.RNA Benchmark Datasets mRNA related tasks Translation efficiency prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split mRNA expression level prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split Mean ribosome load prediction from Sample et al. (2019) [2] input sequence: 5'UTR ouput: mean ribosome load the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.tabular1M<n<10M0 likes1.8k downloads1mo agoHugging Face21Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes1.5k downloads5mo agoHugging Face22joyfine /Qwen3-235B-A22B-Thinking-2507_Qwen3-1.7B_AIME_1983_2024textn<1K0 likes1.3k downloads11mo agoHugging Face23isalgo /airr_controltabular10M<n<100M0 likes1.2k downloads3mo agoHugging Face24cyd0806 /spartina-ai-eco-evolution-data Spartina AI eco-evolutionary reanalysis Code repository: github.com/ydchen0806/spartina-ai-eco-evolution Reproducible code, derived tables and audit figures for the Spartina alterniflora aerial-observation project. The release reconstructs the available 2014–2021 patch-record analysis and audits the recovered 2014–2020 environment–growth simulation archive. Active manuscript and validation pilot The full working manuscript is available as editable DOCX and PDF. It… See the full description on the dataset page: https://huggingface.co/datasets/cyd0806/spartina-ai-eco-evolution-data.imagetabular-regression1K<n<10K0 likes1.1k downloads15h agoHugging Face25maum-ai /CostNav-Teleop-Dataset CostNav Teleop Dataset Dataset Summary The CostNav Teleop Dataset is a large-scale collection of human teleoperation recordings for robot navigation in an urban sidewalk simulation environment. It was collected as part of the CostNav benchmark, which evaluates navigation systems using real-world economic cost and revenue metrics rather than purely technical metrics. The dataset contains 2,203 teleoperation episodes totaling 50.2 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/CostNav-Teleop-Dataset.tabularrobotics1K<n<10K1 likes1.1k downloads4mo agoHugging Face26AIML-TUDA /i2p Inaproppriate Image Prompts (I2P) The I2P benchmark contains real user prompts for generative text2image prompts that are unproportionately likely to produce inappropriate images. I2P was introduced in the 2023 CVPR paper Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. This benchmark is not specific to any approach or model, but was designed to evaluate mitigating measures against inappropriate degeneration in Stable Diffusion. The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/i2p.tabular1K<n<10K21 likes1.1k downloads3y agoHugging Face27infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes963 downloads2y agoHugging Face28NMAIResearch /eu-ai-act-article-50-scoreboard Article 50 historical public-evidence snapshot This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set. Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.imagen<1K0 likes941 downloads18d agoHugging Face29aieng-lab /biasneutral-ajibawa BIASNEUTRAL Ajibawa BIASNEUTRAL Ajibawa contains Ajibawa-derived text snippets that passed the GRADIEND bias-neutral filtering pipeline. It is the Ajibawa-source counterpart to the original aieng-lab/biasneutral dataset, which can not be directly published due to licensing issues. This dataset provides an easier-to-access solution, while maintaining the same generation principles as the original BIASNEUTRAL properties. Usage from datasets import load_dataset ds… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/biasneutral-ajibawa.text1M<n<10M0 likes932 downloads2mo agoHugging Face30shi-labs /physical-ai-bench-conditional-generation Physical AI Bench - Conditional Generation Paper | Code This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below. Dataset Category Sample Nums Agibot World Robotics 200 OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.textvideo-to-videon<1K0 likes874 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.