Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ayushnangia /docmath-eval-failures-200 DocMath-Eval Failures 200: Agent Benchmark & Leaderboard A curated benchmark of 200 challenging financial math questions that leading AI models failed to answer correctly, with comprehensive evaluation results from multiple AI agents. Leaderboard Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring. Rank Agent Model Exact Match Judge: Exact Judge: Approx Judge: Total Wrong Avg Duration Avg Tool Calls 1 TRAE Agent Opus 4.5 98/200 (49.0%) 96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.tabularquestion-answering1K<n<10K0 likes22k downloads8mo agoHugging Face02CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.5k downloads1y agoHugging Face03nvidia /compute-eval Dataset Card for ComputeEval ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.tabulartext-generation1K<n<10K30 likes810 downloads26d agoHugging Face04glayguo /evalarc-independent-swe Independent-source SWE workflow records 36 Qwen3-8B attempts compare four fixed workflows on three public SWE-bench Verified tasks. There are no accepted attempts: 31 have assessable native reports and five retain an upstream infrastructure flag, so their task outcome is uncertain. Eight attempts produced nonempty patches. One generation request has incomplete usage. Inspect the interactive report · English method · 中文方法 · Offline review and exact raw records This dataset… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/evalarc-independent-swe.tabulartext-generationn<1K0 likes486 downloads2d agoHugging Face05dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes351 downloads12d agoHugging Face06YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes324 downloads5mo agoHugging Face07prometheus-eval /peerreview-bench PeerReview Bench CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/ Repository:https://github.com/prometheus-eval/cmu-paper-reviewer Paper:https://arxiv.org/abs/2605.20668 Point of Contact:seungone@kaist.ac.kr Expert-annotated review items from scientific papers, organized for three complementary evaluation tasks. All data in this dataset is intended for evaluation, not training. All configs reference a shared, deduplicated file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.tabulartext-classification10K<n<100K3 likes279 downloads5mo agoHugging Face08ToolGym /long-horizon-eval long-horizon-eval Evaluation results for long-horizon agent performance Dataset Description This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics. Dataset Structure The dataset is organized by model name, with each model having separate JSONL files for different experimental passes. long-horizon-eval/ ├── model-1/ │ ├── pass@1.jsonl │ ├── pass@2.jsonl │ └── pass@3.jsonl ├── model-2/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.tabulartext-generation1K<n<10K0 likes219 downloads9mo agoHugging Face09facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes172 downloads1y agoHugging Face10compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes156 downloads5mo agoHugging Face11CharlieLLL /SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920 SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence. Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.tabulartext-generation1K<n<10K0 likes154 downloads20d agoHugging Face12mcp-tool-shop /jam-rollout-arc-evals Rollout arc — raw generations Every model generation behind the write-ups in mcp-tool-shop-org/ai-jam-sessions under experiments/rollout-arc/p4/. Two things you can do with this. Check our arithmetic. The repo has the readout scripts, the preregistrations and the intervals — but the generations they were computed from are ~51 MB and were never committed, so a clone got the conclusions and no way to recompute them. These are those files, unfiltered. Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.tabulartext-generation1K<n<10K0 likes152 downloads26d agoHugging Face13bbidpa /Rainbow-Pony-100m-Flutter-steps-eval Rainbow-Pony-100M Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.tabulartext-generation1K<n<10K1 likes136 downloads1mo agoHugging Face14LaurelWings /rcga-evaluation-data RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.tabulartext-generation10K<n<100K0 likes134 downloads11d agoHugging Face15build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes121 downloads4mo agoHugging Face16cudabenchmarktest /r8-eval-suite-5bucket ⚠️ CRITICAL: Ollama Inference Flag Required for derived models If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama, you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use. The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag. See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned. R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.tabulartext-generationn<1K0 likes100 downloads6mo agoHugging Face17jang1563 /LabCraft-Eval LabCraft-Eval LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI agents execute benign molecular-microbiology protocols inside a seeded laboratory simulator with task-dependent stochasticity. It pairs task prompts and tool-accessible lab operations with deterministic, multi-axis trajectory scoring. This Hugging Face dataset export is generated from the GitHub repository: https://github.com/jang1563/LabCraft-Eval.git Release Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.texttext-generationn<1K0 likes95 downloads1mo agoHugging Face18bbidpa /Rainbow-Pony-100m-Flutter-direct-eval Rainbow-Pony-100M Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes95 downloads1mo agoHugging Face19beatsprom /stateless-mcp-agent-evaluation-suite-2026 ⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1 ⚡ Overview & Industry Problem As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.tabulartext-generation1K<n<10K0 likes95 downloads18d agoHugging Face20MuseMesh /mume-eval-suites Mume evaluation suites The frozen evaluation sets behind every bits-per-byte number of the Muse Mesh English and Math models (mume-english-125m, mume-math-125m), record for record, including the contamination-clean variants; and training_manifests, the list of source rows each model was trained on (no text), so the training data can be rebuilt from the public sources. From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run… See the full description on the dataset page: https://huggingface.co/datasets/MuseMesh/mume-eval-suites.tabulartext-generation1M<n<10M0 likes94 downloads3d agoHugging Face21deluair /econ-eval econ-eval: how much do you give up by using a cheap model for an economist's work? A reproducible benchmark of frontier and cheap LLMs on the work a trade and policy economist actually does: Balassa RCA from raw BACI values, CAGR and share arithmetic, bank capital and systemic-risk formulas, small trade-data pipeline functions, checking a colleague's numbers, and policy writing in English and Bangla. Every task carries a source field, every reference value is derived from… See the full description on the dataset page: https://huggingface.co/datasets/deluair/econ-eval.tabularquestion-answering1K<n<10K0 likes93 downloads25d agoHugging Face22bbidpa /Qwen2.5-Coder-0.5B-Flutter-direct-eval Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct, a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes85 downloads1mo agoHugging Face23anonymous-noname /econ_eval The Price of Progress: Benchmark-Level LLM Inference Cost Dataset Dataset Summary This dataset combines historical LLM inference prices with benchmark performance scores to construct the largest publicly available benchmark-level LLM price dataset we are aware of. It covers 100+ models across three major benchmarks (GPQA-Diamond, SWE-bench Verified, and AIME) over a two-year window from April 2024 to April 2026, with varying coverage per benchmark. The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-noname/econ_eval.tabulartext-generationn<1K0 likes79 downloads5mo agoHugging Face24bbidpa /Qwen2.5-Coder-0.5B-Flutter-steps-eval Qwen2.5-Coder-0.5B Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps, a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated, until the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps-eval.tabulartext-generation1K<n<10K0 likes77 downloads1mo agoHugging Face25hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes75 downloads24d agoHugging Face26Chess-Nut-Engine /chess-sft-eval Chess SFT Eval & Benchmark Held-out evaluation splits and a frozen benchmark for the Chess SFT training pipeline. Every FEN in these files is excluded from training data via a blocklist to guarantee zero contamination. Eval examples 13,000 Benchmark examples 13,000 Splits 9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate) Format JSONL Training companion Chess-Nut-Engine/chess-sft-data How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.tabulartext-generation10K<n<100K0 likes74 downloads7mo agoHugging Face27evaligo /swe-race SWE-Race (public split) v0.1 95 real concurrency and async-lifecycle bugs from 58 open-source Python projects, packaged as Harbor tasks (the DeepSWE layout) and graded by each project's own hidden tests in a sealed, offline container. A private split of 93 tasks is held out to detect overfitting. Leaderboard, per-task results and every agent trajectory: https://labs.evaligo.com/swe-race Get new results by email: https://labs.evaligo.com/swe-race#signup-form Full task set… See the full description on the dataset page: https://huggingface.co/datasets/evaligo/swe-race.tabulartext-generationn<1K0 likes73 downloads6d agoHugging Face28base-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes70 downloads23d agoHugging Face29caiotheodoro /recon-eval ReconEval — Financial Reconciliation Benchmark Reading results from this benchmark. Four properties of ReconEval shape what a score on it means. Anyone comparing models here should know them. One class can dominate a margin. PARTIAL_MATCH is the highest-variance class between models, and its 32 evaluation items are generated from 9 abbreviation pairs — all of which also appear in the training split, overlap fraction 1.0. On this set, "learned the concept" and "memorised nine… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/recon-eval.tabulartext-generation1K<n<10K0 likes68 downloads1mo agoHugging Face30eval-aware /linuxarena-trajectories LinuxArena Trajectories Full agent trajectories from LinuxBench/LinuxArena evaluations across 14 model/policy combinations and 10 environments. Dataset Description Each row is one complete evaluation trajectory — every tool call the agent made from start to finish, with arguments, outputs, errors, and reasoning. Actions are represented as parallel variable-length lists (one element per action). Two granularity levels are provided per action: Normalized… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/linuxarena-trajectories.tabulartext-generationn<1K1 likes64 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.