Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes21k downloads2d agoHugging Face02LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K71 likes1.7k downloads5mo agoHugging Face03witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face04Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads4d agoHugging Face05YuvrajSingh9886 /bonsai-jetson-benchmark-15w Bonsai Jetson Benchmark — 15W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo Status: Complete — 57 combos (5 models × 12 prompt/gen configs) Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W) Models Model Quant Size Bonsai-1.7B Q1_0 (1-bit) ~237 MB Bonsai-4B Q1_0 (1-bit) ~540 MB Bonsai-8B Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.tabulartext-generationn<1K1 likes1.2k downloads3d agoHugging Face06ktwu01 /benchmark-radar Benchmark Radar Dataset Overview Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115). As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.tabulartext-generation10K<n<100K2 likes648 downloads9h agoHugging Face07OSOmni /os-omni-benchmark OS-Omni Benchmark OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks. Contents data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation. data/tasks.jsonl: JSON Lines copy of the same task index. metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/OSOmni/os-omni-benchmark.imagetext-generationn<1K0 likes340 downloads5mo agoHugging Face08nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes334 downloads24d agoHugging Face09marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes274 downloads11mo agoHugging Face10lars1234 /story_writing_benchmark Story Evaluation Dataset This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks. This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.tabulartext-generation10K<n<100K6 likes254 downloads2y agoHugging Face11abadawi /Cognitive_Atrophy_Benchmark Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.tabulartext-generation10K<n<100K2 likes253 downloads3mo agoHugging Face12omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes218 downloads3mo agoHugging Face13Anonymousbabygoat /auditable-fincrime-benchmark Auditable Financial-Crime Compliance Benchmark A matched-twin benchmark of auditable explanations for financial-crime compliance, spanning two domains — anti-money-laundering (AML) and market abuse. Every scenario is judged not by a yes/no label but against a four-element rubric: the mechanism of the abuse, the rule that is defeated (loophole in AML, breach in market abuse), the distinguishing facts that separate a violation from its innocent twin, and the decision (REPORT /… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousbabygoat/auditable-fincrime-benchmark.tabulartext-classification10K<n<100K0 likes217 downloads5d agoHugging Face14SciCodePile /SciCode-Runnable-Benchmark-Reviewedtabulartext-generationn<1K0 likes210 downloads7mo agoHugging Face15tekosML /qwen36-cross-platform-benchmark Qwen3.6-35B-A3B cross-platform benchmark dataset Complete sanitized evidence for the Qwen3.6 GX10 versus M2 Max benchmark. Headline results At 128K, median cold-prompt TTFT was 45.70 s on GX10 and 552.79 s on M2 Max, a 12.10x difference. Near 256K, GX10 completed and strictly passed 12/12 requests. M2 Max completed 9/12 and strictly passed 7/9 completed requests. On the same Q4_K_M coding control, MTP improved median decode 26.87% on GX10 CUDA and regressed… See the full description on the dataset page: https://huggingface.co/datasets/tekosML/qwen36-cross-platform-benchmark.tabulartext-generationn<1K0 likes194 downloads3mo agoHugging Face16eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes172 downloads1y agoHugging Face17aporthq /vault-benchmark-v1gated APort Vault Benchmark v1 4,371 attacks written by humans in 1,128 sessions of a public competition (March to August 2026), replayed against 14 language models from 8 labs acting as a simulated bank teller (no real money moves), in two conditions: model alone, and behind a pre-action authorization layer. Papers APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport (Paper 2). The replay benchmark and results reported in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/aporthq/vault-benchmark-v1.imagetext-generation100K<n<1M2 likes171 downloads20d agoHugging Face18responsible-ai-labs /indian-responsible-ai-benchmark Indian Responsible AI Benchmark A comprehensive benchmark for evaluating responsible AI behavior in Indian contexts — covering 212 adversarial and safety-critical prompts across 22 evaluation categories, 10 Indian language regions, and 8 Responsible AI dimensions. Why This Benchmark? Most AI safety benchmarks are US/Western-centric. Indian users face unique challenges: Caste dynamics not captured by Western bias benchmarks India/US context confusion (models… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/indian-responsible-ai-benchmark.tabulartext-classificationn<1K1 likes159 downloads4mo agoHugging Face19shunanhe /NPM-Artifact-Explanation-Benchmark NPM-Artifact-Explanation-Benchmark English NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation. This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.tabularimage-to-text10K<n<100K1 likes158 downloads26d agoHugging Face20Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug set 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K1 likes157 downloads5d agoHugging Face21humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes151 downloads4mo agoHugging Face22JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes148 downloads8d agoHugging Face23mayank-dubey-ai /l4-gpu-llm-benchmark-leaderboard 🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB) An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU. 📊 Executive Summary & Key Takeaways ⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face24nlile /math_benchmark_test_saturation LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024) This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems. Original source data: Math Word Problem Solving on MATH (Papers with Code) About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.tabularquestion-answeringn<1K0 likes135 downloads2y agoHugging Face25OiQ /hallucination-autopsy-benchmark Hallucination Autopsy Benchmark A unified, standardized benchmark for cross-model, cross-parameter analysis of LLM hallucination phenomena. Overview This dataset merges multiple hallucination detection benchmarks into a single standardized schema, enabling systematic etiological analysis of why and how different LLM architectures hallucinate under specific configurations. Version: 3.0.0Total Records: 69,002Base Records: 69,002Augmented Records: 0Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/hallucination-autopsy-benchmark.tabularquestion-answering10K<n<100K0 likes134 downloads4mo agoHugging Face26CompilingThings /compile-benchmark CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.tabulartext-generation1K<n<10K1 likes131 downloads27d agoHugging Face27sixstringzen /hemmingway-1-omlx-quantization-benchmark-v1 Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.tabulartext-generationn<1K0 likes129 downloads18d agoHugging Face28himajahealth /qwen3.8-27b-medical-benchmarks Qwen3.8-27B — Medical Benchmark Evaluation Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations (27,312 scored items), under a single fixed inference configuration, with two grading regimes: deterministic and model-graded. The model was evaluated as released. No weights were modified. Field Value Model under evaluation Qwen/Qwen3.8-27B Evaluation dates 2026-09-27 to 2026-09-30 (UTC) Benchmark configurations 32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.tabularquestion-answeringn<1K0 likes129 downloads9d agoHugging Face29axjns /llmfit-benchmarks llmfit Real-World LLM Inference Benchmarks An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit. The initial release contains 1,501 normalized observations: 1,010 unique external-community observations from the repository's 2026-08-10 snapshot. 491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.tabulartext-generation1K<n<10K2 likes128 downloads2mo agoHugging Face30Kicaulah /opencode-ai-benchmark 🌐 OpenCode / Antigravity Protocol (October 2026) Professional AI Engineering & Cybersecurity Benchmark Deep Reasoning • Human-Like Engineering Judgment • Adversarial Traps • Zero Fabrication Live Interactive Leaderboard • Executive Report • 120 Skills Taxonomy • Evaluation Protocol • Quickstart [!WARNING] ⚠️ EXPERIMENTAL TRIAL RELEASE (VERSI UJI COBA) Research Preview Notice: This benchmark dataset, leaderboard, and… See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/opencode-ai-benchmark.tabularquestion-answering1K<n<10K0 likes122 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.