Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes24k downloads5mo agoHugging Face02alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes21k downloads2d agoHugging Face03ai-forever /spellcheck_punctuation_benchmarkRussian Spellcheck Benchmark is a new benchmark for spelling correction in Russian language. It includes four datasets, each of which consists of pairs of sentences in Russian language. Each pair embodies sentence, which may contain spelling errors, and its corresponding correction. Datasets were gathered from various sources and domains including social networks, internet blogs, github commits, medical anamnesis, literature, news, reviews and more.text-generation10K<n<100K5 likes1.8k downloads3y agoHugging Face04LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K71 likes1.7k downloads5mo agoHugging Face05ai-forever /spellcheck_benchmarkRussian Spellcheck Benchmark is a new benchmark for spelling correction in Russian language. It includes four datasets, each of which consists of pairs of sentences in Russian language. Each pair embodies sentence, which may contain spelling errors, and its corresponding correction. Datasets were gathered from various sources and domains including social networks, internet blogs, github commits, medical anamnesis, literature, news, reviews and more.text-generation5 likes1.7k downloads3y agoHugging Face06Open-Style /Open-LLM-Benchmark Open-LLM-Benchmark Dataset Description The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator. License: CC-BY 4.0 Dataset Structure An example of model response files looks as follows: { "question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.texttext-generation100K<n<1M1 likes1.6k downloads2y agoHugging Face07billion-word-benchmark /lm1bA benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.text-generation19 likes1.5k downloads3y agoHugging Face08witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face09YuvrajSingh9886 /bonsai-jetson-benchmark-7w Bonsai Jetson Benchmark: 7W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 7W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt in {256, 512, 1024, 2048} tok x gen in {128, 256, 512} tok · 20 reqs/combo Context: 2560 tok · Concurrency: 1 Part of smolperfleaderboard, a public on-device LLM benchmark leaderboard. Headline metric is output tok/J (tokens per joule), computed over the decode phase. Files Bonsai-*, Ternary-Bonsai-*: per-combo… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-7w.text-generationn<1K0 likes1.5k downloads3d agoHugging Face10Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads4d agoHugging Face11YuvrajSingh9886 /bonsai-jetson-benchmark-25w Bonsai Jetson Benchmark — 25W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 25W Status: Partial — 36 combos Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB — 25W Date: 2026-05-27 22:38Backend: llama.cpp (build-jetson) / CUDA / -ngl 99Platform: NVIDIA Jetson Orin Nano Super 8GB (6-core Cortex-A78AE + Ampere GPU)Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok × 20 reqs/comboKey metric: tok/J = output tokens per second ÷… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-25w.text-generationn<1K0 likes1.3k downloads3d agoHugging Face12YuvrajSingh9886 /bonsai-jetson-benchmark-15w Bonsai Jetson Benchmark — 15W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo Status: Complete — 57 combos (5 models × 12 prompt/gen configs) Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W) Models Model Quant Size Bonsai-1.7B Q1_0 (1-bit) ~237 MB Bonsai-4B Q1_0 (1-bit) ~540 MB Bonsai-8B Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.tabulartext-generationn<1K1 likes1.2k downloads3d agoHugging Face13YuvrajSingh9886 /bonsai-jetson-benchmark-maxn Bonsai Jetson Benchmark — MAXN_SUPER Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: MAXN_SUPER Status: Complete — 57 combos Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB — MAXN_SUPER Date: 2026-05-27 00:02Backend: llama.cpp (build-jetson) / CUDA / -ngl 99Platform: NVIDIA Jetson Orin Nano Super 8GB (6-core Cortex-A78AE + Ampere GPU)Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok × 20 reqs/comboKey metric: tok/J = output… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-maxn.text-generationn<1K1 likes1.1k downloads3d agoHugging Face14RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K13 likes1k downloads2mo agoHugging Face15swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes817 downloads2mo agoHugging Face16obaydata /mcp-agent-trajectory-benchmark MCP Agent Trajectory Benchmark A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces. Designed for training and evaluating tool-use / function-calling capabilities of LLMs. Overview Item Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.texttext-generationn<1K3 likes695 downloads7mo agoHugging Face17HeliumTrades /helium-market-resolution-benchmark What is this? Market Resolution contains 299 ranked option-contract questions. Most test calculations or comparisons from frozen quotes. Sixty test an implied-volatility prior with the premium hidden, and 11 test probability-of-finishing-in-the-money forecasts against a later outcome. It does not measure trading profitability. It measures bounded option-chain reasoning: implied volatility (IV), delta, time value, parity, term structure, relative IV, chain surfaces, and… See the full description on the dataset page: https://huggingface.co/datasets/HeliumTrades/helium-market-resolution-benchmark.texttext-generationn<1K1 likes695 downloads3mo agoHugging Face18ktwu01 /benchmark-radar Benchmark Radar Dataset Overview Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115). As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.tabulartext-generation10K<n<100K2 likes648 downloads12h agoHugging Face19pavelslab-nyu /chess_puzzle_benchmark Chess Puzzle Benchmark Chess puzzle evaluation sets at five difficulty tiers, B1 (easiest) through B5 (hardest). Each example is a chess game given in PGN move notation; the model must produce the next move(s). Two prompt variants The same puzzles are released in two forms that differ only in the prompt suffix: think/ — the prompt ends with a special <T> token. <T> is a reasoning trigger: it tells the model to think (produce a chain of reasoning) before… See the full description on the dataset page: https://huggingface.co/datasets/pavelslab-nyu/chess_puzzle_benchmark.texttext-generation1K<n<10K0 likes644 downloads3mo agoHugging Face20sfd-anonymous /edgar-forecast-benchmark EDGAR-Forecast Benchmark EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings. Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs. Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.text-generation0 likes610 downloads5mo agoHugging Face21zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K61 likes594 downloads1mo agoHugging Face22HeliumTrades /helium-model-worldview-benchmark What is this? Model Worldview is a 323-item probe suite. It includes 93 paired tests that hold a scenario fixed while changing a source, identity, or framing cue, plus standalone value tradeoffs, political survey items, and evidence questions. It is not a single left-right score or a ranking of the "best worldview." Across 16 matched stereotype-essay prompts, 6 models triggered the refusal classifier every time. The lowest rate was grok-4.20-reasoning: 3/16. Same requests… See the full description on the dataset page: https://huggingface.co/datasets/HeliumTrades/helium-model-worldview-benchmark.texttext-generationn<1K1 likes559 downloads3mo agoHugging Face23RegalFire /Agent-Failure-Recovery-Benchmark Agent Failure Recovery Benchmark 22,573 synthetic, source-verified failure → recovery trajectories across four domains. Evaluate whether a system can reject a failed plan, choose a recovery, or recognize that no recovery exists. Every row links to a hashed record in a pinned public source and is checked by independent computational oracles and trajectory replay. Tasks: failure diagnosis, recovery-action prediction, planning regression, scoped agent evaluation and imitation… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Agent-Failure-Recovery-Benchmark.textreinforcement-learning10K<n<100K1 likes461 downloads4d agoHugging Face24microsoft /delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark A verified multilingual benchmark for code-completion hallucinations. Every golden completion compiles. Every hallucination provably doesn't. 📄 Read the preprint on arXiv → Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.texttext-generation1K<n<10K3 likes396 downloads5mo agoHugging Face25EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes376 downloads2mo agoHugging Face26ef-ai /wallet-eval-benchmark Wallet tool-calling eval benchmark The eval side of the wallet fine-tuning work: what the models are scored on. The training rows are deliberately not published. These cases are held out from them by construction, and that is the only reason a score here means anything. If you train on this benchmark, say so — a number from a contaminated run is not comparable to the ones below. The model these cases were used to select is public: ef-dai-team/gemma-4-E4B-wallet-ft-v5, which… See the full description on the dataset page: https://huggingface.co/datasets/ef-ai/wallet-eval-benchmark.text-generation1K<n<10K0 likes346 downloads2mo agoHugging Face27latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K9 likes340 downloads3mo agoHugging Face28OSOmni /os-omni-benchmark OS-Omni Benchmark OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks. Contents data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation. data/tasks.jsonl: JSON Lines copy of the same task index. metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/OSOmni/os-omni-benchmark.imagetext-generationn<1K0 likes340 downloads5mo agoHugging Face29csankalp21 /eulerfold-benchmark EulerFold Technical & Scientific Benchmark Dataset The EulerFold Benchmark is a unified, multi-format academic dataset containing 10,017 verified technical problems mapped directly across 10 foundational domains, 329 subject areas, and 3,030 standardized subtopics. Designed to evaluate complex reasoning and power adaptive technical learning systems, the benchmark spans multiple-choice diagnostics with structured misconception maps, open-ended mathematical derivations and… See the full description on the dataset page: https://huggingface.co/datasets/csankalp21/eulerfold-benchmark.textquestion-answering10K<n<100K0 likes338 downloads11d agoHugging Face30Akanezora /TACO-Benchmark TACO-Benchmark TACO (Text-to-SQL with Ambiguous and Cross-database Open-domain queries) is a benchmark for real-world data-lake Text-to-SQL. 📢 News (2026): TACO has been accepted to VLDB 2026! 🎉📄 Paper: arXiv:2606.14201 GitHub (code & evaluation): Akanezora0/TACO-Benchmark Google Drive mirror: TACO-Benchmark.zip Overview Unlike Spider or BIRD — where the target database is known and schemas are clean — TACO evaluates systems on three challenges common in… See the full description on the dataset page: https://huggingface.co/datasets/Akanezora/TACO-Benchmark.text-generation10K<n<100K1 likes337 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.