Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes21k downloads2d agoHugging Face02YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes5.5k downloads1mo agoHugging Face03dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.4k downloads9d agoHugging Face04ai-humanizer-benchmark /ai-humanizer-benchmark AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026) AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.tabular1K<n<10K2 likes2.7k downloads8d agoHugging Face05drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face06EleutherAI /hack-ignition-benchmark hack-ignition benchmark — data, v0.1.6 Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt, training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.tabular100K<n<1M1 likes1.2k downloads19d agoHugging Face07ktwu01 /benchmark-radar Benchmark Radar Dataset Overview Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115). As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.tabulartext-generation10K<n<100K2 likes648 downloads3h agoHugging Face08Keh0t0 /scene-mem-benchmark scene-mem-benchmark A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work in a house for several minutes, then is asked to retrieve an object it has to remember — one that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the job it was interrupted in, remembering how far it had got — or (routine) to put a new object away where this household keeps that kind of thing, a rule it was never told and can only… See the full description on the dataset page: https://huggingface.co/datasets/Keh0t0/scene-mem-benchmark.tabularrobotics1K<n<10K0 likes507 downloads23d agoHugging Face09ibm-research /data-product-benchmark DPDisc Dataset Paper | Code Dataset Description This dataset provides a benchmark for automatic data product creation. The task is framed as follows: given a natural language data product request and a corpus of text and tables, the objective is to identify the relevant tables and text documents that should be included in the resulting data product which would useful to the given data product request. The benchmark brings together three variants: HybridQA, TAT-QA, and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/data-product-benchmark.texttable-question-answering10K<n<100K3 likes399 downloads7mo agoHugging Face10EstherrrCheng /mea-benchmark MEA-Benchmark Benchmark for MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations. 📄 Paper: arXiv:2610.02480 💻 Code: github.com/AikyamLab/xai-agent MEA-Benchmark evaluates explanations of neural network models across three modalities (tabular, text, vision) with ten question types (Q1–Q10), spanning feature attribution, counterfactual reasoning, and spurious feature detection. Each question type is paired with a perturbation-based faithfulness metric (see… See the full description on the dataset page: https://huggingface.co/datasets/EstherrrCheng/mea-benchmark.textquestion-answering10K<n<100K0 likes379 downloads3d agoHugging Face11Jiazuo98 /Finers-4k-benchmarkimage10K<n<100K0 likes276 downloads10mo agoHugging Face12marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes274 downloads11mo agoHugging Face13llmware /rag_instruct_benchmark_tester Dataset Card for RAG-Instruct-Benchmark-Tester Dataset Summary This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases, contracts, invoices, technical articles, general news and short texts. The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.tabularn<1K59 likes263 downloads3y agoHugging Face14TrialPanorama /TrialPanorama-benchmarkDataset website: https://ryanwangzf.github.io/projects/trialpanorama tabular100K<n<1M5 likes243 downloads1y agoHugging Face15utahnlp /knows-benchmark KNOWS Benchmark KNOWS evaluates web agents on the work people actually do in Google Workspace: writing documents, building spreadsheets, and composing slide decks that require web research, multi-step tool use, and faithful grounding in retrieved sources. This dataset contains the task definitions — the prompt an agent receives, plus the structured evaluation rubric used to grade the artifact it produces. Tasks 110 (22 templates × 5 instances) Domains 20 Mean… See the full description on the dataset page: https://huggingface.co/datasets/utahnlp/knows-benchmark.tabularothern<1K3 likes243 downloads10d agoHugging Face16gemmozero /ai-benchmarks-v2-2026gated Ai Benchmarks V2 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-benchmarks-v2-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-benchmarks-v2-2026.tabulartext-classificationn<1K0 likes238 downloads10d agoHugging Face17SciCodePile /SciCode-Runnable-Benchmark-Reviewedtabulartext-generationn<1K0 likes210 downloads7mo agoHugging Face18vkatg /streaming-phi-deidentification-benchmark Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.documentn<1K0 likes202 downloads7mo agoHugging Face19dsadd4 /UltraMS-Benchmark-Assets UltraMS benchmark assets Inputs for reproducing the UltraMS spectral-property and molecular-identification benchmarks. The GitHub benchmark guides provide the training and evaluation commands. Path Use spectral_properties/masked_peak_reconstruction.pt UltraMS checkpoint for masked peak reconstruction. spectral_properties/massspecgym_labels.parquet Frozen MassSpecGym property and neutral-loss training table used in Figure 2. spectral_properties/heteroatom_count/… See the full description on the dataset page: https://huggingface.co/datasets/dsadd4/UltraMS-Benchmark-Assets.tabular100K<n<1M0 likes196 downloads15d agoHugging Face20eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes172 downloads1y agoHugging Face21as-benchmark-artifacts /vqa-cmsv-benchmark VQA-CMSV Benchmark Data Package This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments. Contents data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl masks/vqa_v2_cmsv_masks.npz masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.tabularvisual-question-answering10K<n<100K0 likes143 downloads5mo agoHugging Face22anonymous-ed-benchmark /SKILLRET SkillRet Benchmark SkillRet is a retrieval benchmark for matching natural-language user requests to agent skills. Each retrieval document is a full agent skill, represented by its name, short description, and full Markdown skill body. Each query describes a realistic user request that requires one or more relevant skills. The benchmark is built from public agent skills indexed from GitHub and contains synthetic train and evaluation queries generated through a self-instruct-style… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-ed-benchmark/SKILLRET.tabulartext-retrieval100K<n<1M0 likes134 downloads15d agoHugging Face23meme-benchmark /MEME MEME: Multi-Entity and Evolving Memory Evaluation A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.tabularquestion-answeringn<1K5 likes133 downloads5mo agoHugging Face24shounakpaul95 /Benchmark-Testingtabulartext-classification100K<n<1M0 likes132 downloads2y agoHugging Face25CompilingThings /compile-benchmark CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.tabulartext-generation1K<n<10K1 likes131 downloads27d agoHugging Face26mcphunt-benchmark /mcphunt-agent-traces MCPHunt Agent Traces Agent execution traces from the MCPHunt evaluation framework, measuring cross-boundary data propagation in multi-server MCP agents. Contents main/ — 3,615 traces from 5 models across 147 tasks and 7 environment variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model. mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3 levels) across 3 models. live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/mcphunt-benchmark/mcphunt-agent-traces.tabularothern<1K3 likes129 downloads2mo agoHugging Face27sixstringzen /hemmingway-1-omlx-quantization-benchmark-v1 Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.tabulartext-generationn<1K0 likes129 downloads18d agoHugging Face28retroam /repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes124 downloads3mo agoHugging Face29lips-poc /powergrid-benchmark2 Note: This is a dummy dataset for browsing only. The official dataset can be downloaded from our pipeline (Data Hub). tabularn<1K0 likes124 downloads11d agoHugging Face30rayrren /peptide-reasoning-benchmark Peptide Reasoning Benchmark PEB v1.0-RC benchmark release for peptide-reasoning model evaluation. Includes cases, splits, baselines, references, and leaderboard artifacts. GitHub: https://github.com/ray-r-ren/peptide-reasoning-bench Trained a small reference LoRA model: https://huggingface.co/rayrren/the-spice-v0-mvp tabular1K<n<10K0 likes121 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.