Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes22k downloads5mo agoHugging Face02alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes19k downloads18h agoHugging Face03Open-Style /Open-LLM-Benchmark Open-LLM-Benchmark Dataset Description The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator. License: CC-BY 4.0 Dataset Structure An example of model response files looks as follows: { "question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.texttext-generation100K<n<1M1 likes1.1k downloads2y agoHugging Face04RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K12 likes997 downloads2mo agoHugging Face05swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes965 downloads2mo agoHugging Face06zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K60 likes878 downloads1mo agoHugging Face07obaydata /mcp-agent-trajectory-benchmark MCP Agent Trajectory Benchmark A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces. Designed for training and evaluating tool-use / function-calling capabilities of LLMs. Overview Item Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.texttext-generationn<1K3 likes724 downloads6mo agoHugging Face08ktwu01 /benchmark-radar Benchmark Radar Dataset Overview Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115). As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.tabulartext-generation10K<n<100K2 likes548 downloads16h agoHugging Face09EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes378 downloads1mo agoHugging Face10TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes362 downloads1y agoHugging Face11Parakeet-Inc /joyo-kanji-yomi-benchmark-parakeet 日本語 | English 常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet) 常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。 このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。 このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。 概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.texttext-to-speech10K<n<100K5 likes349 downloads2mo agoHugging Face12marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes267 downloads11mo agoHugging Face13springofwindslabs /mcp-agent-trajectory-benchmark 🛑 Stop LLM Agent Collapse Caused by State Corruption Most agent failures are not tool failures. They are state failures. The model believes the world is still valid — when reality has already changed. This MCP trajectory dataset trains belief revision, state recovery, and autonomous replanning: believed : 45/50 rooms synced ← agent trusts a stale worker log tool : "success", count: 0 ← silent no-op (the worst failure mode) verify : GDS reports actual = 40 ←… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.texttext-generationn<1K0 likes260 downloads13h agoHugging Face14llmsql-bench /llmsql-benchmark LLMSQL Benchmark ⚠️ A newer version of this dataset is available:👉 https://huggingface.co/datasets/llmsql-bench/llmsql-2.0 This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark. Arxiv Article: https://arxiv.org/abs/2510.02350 Files tables.jsonl — Database table metadata questions.jsonl — All available questions train_questions.jsonl, val_questions.jsonl, test_questions.jsonl — Data… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-benchmark.textquestion-answering10K<n<100K2 likes217 downloads7mo agoHugging Face15sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes201 downloads5mo agoHugging Face16idleengine /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/big-finance-benchmark.textquestion-answeringn<1K0 likes194 downloads1mo agoHugging Face17adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes190 downloads4mo agoHugging Face18SciCodePile /SciCode-Runnable-Benchmark-Reviewedtabulartext-generationn<1K0 likes186 downloads7mo agoHugging Face19oliversayshi /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.textquestion-answeringn<1K0 likes183 downloads3mo agoHugging Face20CompilingThings /compile-benchmark CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.tabulartext-generation1K<n<10K1 likes180 downloads23d agoHugging Face21hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes179 downloads13d agoHugging Face22Linmumu009 /LogiTraj-Benchmark LogiTraj Benchmark Dual license. Dataset material is licensed under CC BY 4.0; software in evaluation/code/ is licensed under Apache-2.0. See LICENSE, LICENSE-DATA, and LICENSE-CODE. Commercial-model raw outputs are not included. Synthetic Chinese enterprise logistics sandboxes, tasks, documents, versioned evaluators, verdicts, and Core/Silver/Audit quality views. Task-family coverage is source-faithful rather than imputed: the 20260628_v45, 20260628_v46, and 20260628_v50 SFT… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Benchmark.textquestion-answering10K<n<100K0 likes162 downloads2mo agoHugging Face23eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes160 downloads1y agoHugging Face24lawful-good-project /sud-resh-benchmark Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права. Бенчмарк создан на вычислительный грант от Yandex OpenSource Бенчмарк основан на анонимизированных решениях судов в следующих отраслях: Административное право Конституционное право Экологическое право Финансовое право Гражданское право Семейное право Право социального обеспечения Трудовое право Уголовное право Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.texttext-generation1K<n<10K4 likes156 downloads1y agoHugging Face25sharkiefff /RBAC-Text2SQL-Benchmark RBAC-Text2SQL Benchmark Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural language question with a role policy; the model must either produce a correct SQL query that touches only authorized resources, or refuse with Sorry, I cannot answer. Code, evaluation harness, and reproduction instructions: https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.texttext-generation10K<n<100K1 likes143 downloads2mo agoHugging Face26UltraRAG /UltraRAG_Benchmark UltraRAG 2.0: Accelerating RAG for Scientific Research UltraRAG 2.0 (UR-2.0) is jointly released by THUNLP, NEUIR, OpenBMB, and AI9Stars. It is the first lightweight RAG system construction framework built on the Model Context Protocol (MCP) architecture, designed to provide efficient modeling support for scientific research and exploration. The framework offers a full suite of teaching examples from beginner to advanced levels, integrates 17 mainstream benchmark tasks and a wide… See the full description on the dataset page: https://huggingface.co/datasets/UltraRAG/UltraRAG_Benchmark.textquestion-answering100K<n<1M6 likes128 downloads1y agoHugging Face27batuhanaktas /kids-multilingual-benchmark TinyAya v2 — Multilingual Benchmark for Children's AI Companions 2,312 child–AI conversational prompts across 23 languages, evaluated against four models with five-judge LLM-as-judge validation. 📄 Companion article: see HF Articles by @batuhanaktas. 💻 Code: https://github.com/aktasbatuhan/cohere-tiny-aya-for-kids Dataset summary This dataset contains: benchmark/items.jsonl — 2,312 benchmark items in 23 languages. Each item is a structured prompt designed to mimic… See the full description on the dataset page: https://huggingface.co/datasets/batuhanaktas/kids-multilingual-benchmark.texttext-generation10K<n<100K1 likes128 downloads5mo agoHugging Face28sixstringzen /hemmingway-1-omlx-quantization-benchmark-v1 Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.tabulartext-generationn<1K0 likes125 downloads13d agoHugging Face29Koplos /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/big-finance-benchmark.textquestion-answeringn<1K0 likes123 downloads2mo agoHugging Face30YouMindInc /agentic-task-benchmark YouMind Agentic Task Benchmark v0.1.1 · Experimental YouMindInc/agentic-task-benchmark is a small experimental benchmark for evaluating creative and research task outcomes against selected references. This release contains four task descriptions, a standard outcome format, and an offline text scorer. Task definitions and evaluation protocols may change. Tasks Configuration Task Status image_generation Riso portrait Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.texttext-generationn<1K0 likes117 downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.