Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes21k downloads2d agoHugging Face02leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes9.8k downloads6mo agoHugging Face03madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M1 likes8.7k downloads1mo agoHugging Face04YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes5.5k downloads1mo agoHugging Face05gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes4.3k downloads5h agoHugging Face06inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes4.3k downloads4mo agoHugging Face07BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face08dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.4k downloads9d agoHugging Face09NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Package version recorded in run metadata: 0.2.0 (some runs lack version metadata). Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes3.2k downloads13d agoHugging Face10NoeFlandre /geoparser-benchmark-results Geoparser benchmark results Geoparsing pipelines from geoparser scored on English and multilingual corpora, run on Grid'5000 (one Tesla T4). Benchmarks Benchmark Languages Docs Toponyms Source GeoVirus en 229 2167 WikiNews articles on epidemics (Gritta et al., 2018). HIPE-2020 de, en, fr 129 1516 Historical Swiss, Luxembourgish and American newspapers, OCR. NewsEye de, fi, fr, sv 77 1772 Historical European newspapers, OCR (HIPE-2022). TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.tabularn<1K0 likes2.9k downloads6d agoHugging Face11ai-humanizer-benchmark /ai-humanizer-benchmark AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026) AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.tabular1K<n<10K2 likes2.7k downloads9d agoHugging Face12brettsp /stan-benchmarktabular1M<n<10M0 likes2.4k downloads8h agoHugging Face13drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face14inria-soda /tabular-benchmark Tabular Benchmark Dataset Description This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms. Repository: https://github.com/LeoGrin/tabular-benchmark/community Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document Dataset Summary Benchmark made of curation of various tabular data learning tasks, including: Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.tabulartabular-classification10M<n<100M51 likes2.2k downloads3y agoHugging Face15brennercruvinel /mtg-urna-benchmark 38,627 Magic: The Gathering cards, one per oracle id, the scan and the rules text of each, packed into single .urna files that answer text and image queries from memory-mapped bytes: no server, no Python at read time. Ten such files live here. They carry the same cards, the same text and the same content hash; what differs is how the 4 GB of JPEG was encoded inside, and which image models embedded it. The file to start with is release/v0.3/stills-5models. It is the only one with the models… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/mtg-urna-benchmark.imageimage-to-text100K<n<1M2 likes2k downloads5d agoHugging Face16nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for 8 language subsets, with 10,100 examples selected from Common Voice 17.0. Each example supplies a speaker reference and an independently selected target text, with human recordings as WER/CER and speaker-similarity anchors when the corresponding audio is available. Corpus WER and the WavLM-FT evaluation follow seed-tts-eval. Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes1.8k downloads4d agoHugging Face17physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes1.8k downloads4mo agoHugging Face18LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K71 likes1.7k downloads5mo agoHugging Face19lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.6k downloads3mo agoHugging Face20x-square-robot /xplanner-benchmark XPlanner-Benchmark XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner. Contents 1,500 episodes 3,490 MP4 video references 167 source dataset identifiers 525 unique task names 31 task classes 41 inferred atomic action labels 22.70 total hours of episode… See the full description on the dataset page: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark.tabularrobotics1K<n<10K2 likes1.6k downloads17d agoHugging Face21witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face22Merserk /Krea-2-Turbo-Checkpoint-Format-Benchmark Krea 2 Turbo ComfyUI Format Fidelity Benchmark This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code. Main result BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.imagetext-to-imagen<1K5 likes1.5k downloads3mo agoHugging Face23Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads3d agoHugging Face24YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-15w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99 Tok/s Req/s E2E avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.tabularn<1K0 likes1.3k downloads3d agoHugging Face25Poupou /oparq-benchmarks oparq benchmark results source code · PyPI package · benchmark inputs Results only: 21 datasets, 967,938,981 input rows, 986 source files. No original source rows are distributed. The viewer rows describe algorithm measurements, not individual source events. Full-corpus results sort all input rows globally using fixed writer settings; only key planning is sampled. The separate Arrow/DuckDB comparison sorts within each physical file using saved keys, and must not be conflated… See the full description on the dataset page: https://huggingface.co/datasets/Poupou/oparq-benchmarks.tabularn<1K0 likes1.3k downloads4d agoHugging Face26YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-7w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99 Tok/s Req/s E2E avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.tabularn<1K0 likes1.2k downloads3d agoHugging Face27OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-10-10 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.2k downloads14h agoHugging Face28YuvrajSingh9886 /bonsai-jetson-benchmark-15w Bonsai Jetson Benchmark — 15W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo Status: Complete — 57 combos (5 models × 12 prompt/gen configs) Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W) Models Model Quant Size Bonsai-1.7B Q1_0 (1-bit) ~237 MB Bonsai-4B Q1_0 (1-bit) ~540 MB Bonsai-8B Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.tabulartext-generationn<1K1 likes1.2k downloads3d agoHugging Face29YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-25w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.tabularn<1K0 likes1.2k downloads3d agoHugging Face30YuvrajSingh9886 /jetson-non-reasoning-benchmark-7w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3) Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w Note: tok/J computed from per-run start_time/end_time in each aiperf JSON Full Results Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.tabularn<1K0 likes1.2k downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.