Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-benchmarks /transformers2 likes27k downloads7h agoHugging Face02MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K49 likes13k downloads2mo agoHugging Face03leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes12k downloads6mo agoHugging Face04madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M0 likes8.8k downloads27d agoHugging Face05miromind-ai /MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker 7 likes3.8k downloads9mo agoHugging Face06RedHatAI /speculator_benchmarksThis dataset contains dataset splits for evaluating speculative decoding algorithms on different tasks. File: Coding: HumanEval.jsonl Math: math_reasoning.jsonl Question Answering: qa.jsonl MT_bench: question.jsonl Retrieval-Augmented Generation: rag.jsonl Summarization: summarization.jsonl Translation (German to English): translation.jsonl Writing: writing.jsonl The data comes from two sources: https://github.com/openai/human-eval (1). (The MIT License)… See the full description on the dataset page: https://huggingface.co/datasets/RedHatAI/speculator_benchmarks.4 likes3.3k downloads6mo agoHugging Face07LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes2.7k downloads1y agoHugging Face08aurelio-amerio /SBI-benchmarkstimeseries1M<n<10M1 likes2k downloads2mo agoHugging Face09oking0197 /graphmemix-benchmarks GraphMemix Benchmarks Unified multimodal memory benchmark bundles used by GraphMemix (arXiv:2608.26983) — four long-term personalized memory benchmarks with their raw media assets, packaged together for reproducible evaluation. Benchmark Questions Memories Track Upstream license ATM-Bench (default + hard) 1,044 11,034 memory QA over one multimodal archive MIT Mem-Gallery 1,711 7,944 multimodal gallery memory QA MIT MemEye 1,855 3,392 comics-derived memory QA… See the full description on the dataset page: https://huggingface.co/datasets/oking0197/graphmemix-benchmarks.image1K<n<10K0 likes1.8k downloads1mo agoHugging Face10gtysssp /audio_benchmarksaudion<1K2 likes1.7k downloads1y agoHugging Face11witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads16d agoHugging Face12OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-10-06 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.3k downloads2h agoHugging Face13scaledown /vllm-inference-benchmarkstextn<1K0 likes1.1k downloads8d agoHugging Face14satissss /Squrve-Benchmarks0 likes828 downloads8mo agoHugging Face15yyyang /UI-Grounding-Benchmarks UI-Grounding-Benchmarks This is a collection of UI grounding benchmarks: ScreenSpot ScreenSpot-V2 ScreenSpot-Pro OS-World-G UI-Vision Thanks for their great work! This benchmark collection is used in the paper: FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection 🖼️ Project Page: https://showlab.github.io/FocusUI/ 🏠 Github Repo: https://github.com/showlab/FocusUI 📝 Paper: https://arxiv.org/pdf/2601.03928 Model Zoo Model Backbone 🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.image1K<n<10K2 likes827 downloads8mo agoHugging Face16LLDDSS /Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchimage1K<n<10K0 likes778 downloads1y agoHugging Face17jinulee-v /expert-rag-benchmarks Expert RAG Benchmarks A unified collection of four expert-level legal RAG benchmarks, exposed as six named splits and three relational configurations: questions, documents, and qrels. The KCL split is named kcl_essay because Hugging Face split identifiers do not permit hyphens; its source name remains kcl-essay. Loading from datasets import load_dataset repo_id = "jinulee-v/expert-rag-benchmarks" questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.textquestion-answering1M<n<10M0 likes770 downloads4d agoHugging Face18generative-graphics /genvsr-video-benchmarksimage1K<n<10K0 likes709 downloads8d agoHugging Face19chn123 /spectre-ctrate-bimcv-500-benchmarks v5 correction: SPECTRE embeddings re-extracted with correct resampling The SPECTRE embeddings released in v4 (spectre_ctrate_embeddings.npz, spectre_bimcv_embeddings.npz, their *_500_* copies, ctrate_results/, bimcv_results/ and all SPECTRE metrics derived from them) are invalid and should not be used. Cause: spectre-fm 0.2.1 spectre.io.resample() passes the voxel spacing to MONAI Spacing through the deprecated affine= argument, which MONAI >= 0.9 ignores. Every scan was… See the full description on the dataset page: https://huggingface.co/datasets/chn123/spectre-ctrate-bimcv-500-benchmarks.0 likes608 downloads20d agoHugging Face20katarinagresova /Genomic_Benchmarks_human_enhancers_cohn Dataset Card for "Genomic_Benchmarks_human_enhancers_cohn" More Information needed text10K<n<100K2 likes584 downloads4y agoHugging Face21fastbuilderai /fastmemory-supremacy-benchmarks FastMemory: Beyond A Million (BEAM) 10M Audit Auditing Architectural Integrity at Scale (30 SOTA Wins) This repository contains the official evaluation logs, simulation code, and technical whitepapers for FastMemory’s 10 Million Token BEAM Benchmark Study. FastMemory is a sovereign, local-first memory architecture for agentic AI. Unlike traditional vector-based RAG, FastMemory utilizes Topological Isolation to achieve 100% precision in mission-critical reasoning tasks across… See the full description on the dataset page: https://huggingface.co/datasets/fastbuilderai/fastmemory-supremacy-benchmarks.1 likes569 downloads6mo agoHugging Face22gitpullpull /D-CoT-Benchmarks0 likes525 downloads8mo agoHugging Face23deepinv /benchmarks0 likes518 downloads1mo agoHugging Face24BitRouterAI /benchmarks BitRouter Benchmarks This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/. Main result All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.tabular100K<n<1M1 likes516 downloads27d agoHugging Face25monteirot /lra-benchmarks10B<n<100B0 likes496 downloads8mo agoHugging Face26MothMalone /data-preprocessing-automl-benchmarks Data Preprocessing AutoML Benchmarks This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML. Usage Load a specific dataset configuration like this: from datasets import load_dataset # Example for loading the TREC dataset dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec") Available Datasets Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.texttext-classification100K<n<1M0 likes483 downloads1y agoHugging Face27katarinagresova /Genomic_Benchmarks_human_nontata_promoters Dataset Card for "Genomic_Benchmarks_human_nontata_promoters" More Information needed text10K<n<100K0 likes446 downloads4y agoHugging Face28lthn /LEM-benchmarks LEM-benchmarks Canonical 8-PAC benchmark results for the Lemma model family. This dataset is an aggregated store of per-round evaluation data produced by lthn/LEM-Eval. Every row represents one model's answer to one question in one round of a paired A/B run against its unmodified base, and the dataset grows monotonically as more workers contribute — different machines, different sampling states, different hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.tabularquestion-answering10K<n<100K3 likes443 downloads6mo agoHugging Face29katarinagresova /Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs" More Information needed text100K<n<1M4 likes405 downloads3y agoHugging Face30chn123 /mg3d-ctrate-bimcv-500-benchmarks MG-3D Swin-B (MIA 2026) Benchmark on CT-RATE 500 & BIMCV-R 500 Cohorts (v4 Definitive Release) This repository contains the definitive out-of-fold benchmark results, extracted 3D embeddings, probability predictions, affine/HU calibration manifests, and reproducibility code for MG-3D Swin-B (MIA 2026) evaluated across two independent clinical 3D Computed Tomography cohorts: CT-RATE Validation Cohort (500 series, 428 unique patients, 463 unique studies) BIMCV-R Cohort (317… See the full description on the dataset page: https://huggingface.co/datasets/chn123/mg3d-ctrate-bimcv-500-benchmarks.0 likes385 downloads20d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.