Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlphaDojo /dojo_benchmark_kline Languages: 简体中文 · English dojo_benchmark_kline — Benchmark Index Bars Overview Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date. Files File Description data.parquet Index daily bars Key Fields Field Description symbol Index code (e.g. ^SPX, 000300.SS) kline_t Bar interval; "1D" for daily bars bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.text1K<n<10K1 likes16k downloads4h agoHugging Face02BLINK-Benchmark /BLINK BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.image1K<n<10K49 likes14k downloads1y agoHugging Face03MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K49 likes13k downloads2mo agoHugging Face04leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes12k downloads6mo agoHugging Face05gaia-benchmark /GAIAgated GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.audion<1K865 likes11k downloads11mo agoHugging Face06Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M36 likes9.8k downloads1mo agoHugging Face07nlile /hendrycks-MATH-benchmark Hendrycks MATH Dataset Dataset Description The MATH dataset is a collection of mathematics competition problems designed to evaluate mathematical reasoning and problem-solving capabilities in computational systems. Containing 12,500 high school competition-level mathematics problems, this dataset is notable for including detailed step-by-step solutions alongside each problem. Dataset Summary The dataset consists of mathematics problems spanning multiple… See the full description on the dataset page: https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark.text10K<n<100K33 likes9.3k downloads2y agoHugging Face08dpdl-benchmark /oxford_flowers102image1K<n<10K12 likes6.8k downloads2y agoHugging Face09md-nishat-008 /mHumanEval-Benchmark 🔷 Accepted in NAACL Proceedings (2025) 🔷 mHumanEval The mHumanEval benchmark is curated based on prompts from the original HumanEval 📚 [Chen et al., 2021]. It includes a total of 33,456 prompts for Python, and 836,400 in total - significantly expanding from the original 164. Quick Start Detailed… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/mHumanEval-Benchmark.text100K<n<1M4 likes6.1k downloads1y agoHugging Face10BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes5.8k downloads1y agoHugging Face11inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes5.6k downloads4mo agoHugging Face12gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes4.1k downloads2h agoHugging Face13tccoin /navverse-benchmark NavVerse Benchmark This repository hosts the NavVerse dataset. Runnable scene release navverse_v1_july23.tar.gz is the current runnable scene package. It contains: 52 VC+ outdoor scenes and 52 connected indoor/outdoor scenes under vc_plus/; 30 GRScenes commercial navigation scenes under grscenes_commercial/; the corresponding prebuilt navmesh/ assets; a relative nvidia -> vc_plus/nvidia link for the bundled CloudySky runtime lighting. From a NavVerse-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/tccoin/navverse-benchmark.imagen<1K1 likes3.5k downloads2mo agoHugging Face14futurehouse /ether0-benchmark ether0-benchmark QA benchmark (test set) for the ether0 reasoning language model: https://huggingface.co/futurehouse/ether0 This benchmark is made from commonly used tasks - like reaction prediction in USPTO/ORD, molecular captioning from PubChem, or predicting GHS classification. It's unique from other benchmarks in that all answers are a molecule. It's balanced so that each task is about 25 questions, a reasonable amount for frontier model evaluations. The tasks generally follow… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/ether0-benchmark.textquestion-answeringn<1K17 likes3.5k downloads1y agoHugging Face15LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes2.7k downloads1y agoHugging Face16OpenSTEF /liander2024-energy-forecasting-benchmark Dataset Card for Liander 2024 Short Term Energy Forecasting Benchmark This dataset provides a benchmark for short term energy forecasting models, combining electrical load measurements from Dutch DSO Liander with predictors like corresponding weather data from OpenMeteo, day-ahead electricity prices from ENTSO-E, and profiles of electricity consumption from Energiedatawijzer. The dataset covers the full year 2024 (2024-01-01 to 2025-01-01 UTC) and includes 55 different points in… See the full description on the dataset page: https://huggingface.co/datasets/OpenSTEF/liander2024-energy-forecasting-benchmark.time-series-forecasting10M<n<100M5 likes2.6k downloads11mo agoHugging Face17Avature /MELO-Benchmark MELO Benchmark This dataset contains the Multilingual Entity Linking of Occupations (MELO) Benchmark for easy loading with the HuggingFace datasets library. Dataset Description MELO is a collection of 48 datasets for evaluating the linking of entity mentions in 21 languages to the ESCO Occupations multilingual taxonomy. It was built using high-quality, pre-existent human annotations. Original paper abstract: We present the Multilingual Entity Linking of Occupations… See the full description on the dataset page: https://huggingface.co/datasets/Avature/MELO-Benchmark.texttext-retrieval1M<n<10M0 likes2.6k downloads9mo agoHugging Face18CritPt-Benchmark /CritPt Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark |🌐 Website | GitHub | 📖 Paper | Dataset description CritPt (Complex Research using Integrated Thinking – Physics Test; reads as "critical point") is the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics… See the full description on the dataset page: https://huggingface.co/datasets/CritPt-Benchmark/CritPt.textn<1K27 likes2.5k downloads11mo agoHugging Face19brettsp /stan-benchmarktabular1M<n<10M0 likes2.4k downloads1h agoHugging Face20FanqingM /MMIU-Benchmark Dataset Card for MMIU Repository: https://github.com/OpenGVLab/MMIU Paper: https://arxiv.org/abs/2408.02718 Project Page: https://mmiu-bench.github.io/ Point of Contact: Fanqing Meng Introduction MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular MLLMs, including both open-source and proprietary models… See the full description on the dataset page: https://huggingface.co/datasets/FanqingM/MMIU-Benchmark.image10K<n<100K11 likes2.2k downloads2y agoHugging Face21Marqo /benchmark-embeddings Marqo Benchmark Embeddings This dataset contains a large collection of embeddings from popular models on benchmark datasets. In addition to this, the passage data also includes the local intrinsic dimensionality (LID) for every vector considering its exact nearest 100 neighbours, LID is calculated using a Maximum Likelihood Estimation based approach. Below is a list of the datasets and the models, every datasets queries and passages are embedded with every model.… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/benchmark-embeddings.text100M<n<1B6 likes2.1k downloads2y agoHugging Face22ade-benchmark-corpus /ade_corpus_v2 Dataset Card for Adverse Drug Reaction Data v2 Dataset Summary ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for Classification if a sentence is ADE-related (True) or not (False) and Relation Extraction between Adverse Drug Event and Drug. DRUG-AE.rel provides relations between drugs and adverse effects. DRUG-DOSE.rel provides relations between drugs and dosages. ADE-NEG.txt provides all sentences in the ADE corpus that DO NOT contain any… See the full description on the dataset page: https://huggingface.co/datasets/ade-benchmark-corpus/ade_corpus_v2.texttext-classification10K<n<100K36 likes2.1k downloads3y agoHugging Face23aurelio-amerio /SBI-benchmarkstimeseries1M<n<10M1 likes2k downloads2mo agoHugging Face24infgrad /PosIR-Benchmark-v1text100K<n<1M3 likes1.9k downloads10mo agoHugging Face25OALL /AlGhafa-Arabic-LLM-Benchmark-Native AlGhafa Arabic LLM Benchmark New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility. Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks: Belebele Ar MSA Bandarkar et al. (2023): 900 entries Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.text10K<n<100K8 likes1.8k downloads3y agoHugging Face26LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K49 likes1.7k downloads5mo agoHugging Face27lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.7k downloads3mo agoHugging Face28physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes1.7k downloads4mo agoHugging Face29lmms-lab-encoder /MMT-Benchmarkimage10K<n<100K0 likes1.6k downloads2y agoHugging Face30mundo-ai /turn-benchmark-devgated TurnBench - Dev Set TurnBench is a benchmark for evaluating conversational turn-taking: end-of-turn and interruption detection on real annotated two-speaker conversations. This repository contains the development split: 38 English conversations, about 7.3 hours of audio, packaged as one row per conversation. Each row contains two time-aligned per-speaker audio streams plus three independent annotator tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.audiovoice-activity-detectionn<1K8 likes1.6k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.