Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uclanlp /OpenVLHarness-Evaluation-Datasets OpenVLHarness evaluation datasets Processed evaluation splits used by OpenVLHarness (project page). Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.imagevisual-question-answering10K<n<100K1 likes1.6k downloads20h agoHugging Face02CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes660 downloads5mo agoHugging Face03RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes420 downloads4y agoHugging Face04egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face05RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes217 downloads4y agoHugging Face06OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes206 downloads3y agoHugging Face07JosephAA /toporisk-evaluation-artifacts TopoRisk Evaluation Artifacts English | 简体中文 TopoRisk studies how a limited verification budget should be allocated across an agentic workflow graph. Instead of ranking steps only by their local failure probability, the scheduler estimates how an error can reach terminal outputs and recomputes marginal value after every selected audit. This repository is an artifact-first research preview. It contains code, aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.tabularn<1K1 likes155 downloads2d agoHugging Face08CrimsonRobot /crimson-ma2-evaluation Crimson MolmoAct2: preliminary real-robot evaluation Right-arm bottle pick-and-place rollouts recorded on 2026-09-24 (Asia/Bangkok). This release contains three checkpoints, 31 logged attempts, videos from three cameras, selected recorded telemetry, and joint trajectories. It is an evaluation evidence release, not a training dataset or a completed six-checkpoint benchmark. Model source: Kavin60606/crimson-ma2-ckpts. The policy task text was pick and place the bottle. Each… See the full description on the dataset page: https://huggingface.co/datasets/CrimsonRobot/crimson-ma2-evaluation.tabularroboticsn<1K0 likes92 downloads16d agoHugging Face09krishagarwal /evaluation5tabularn<1K0 likes86 downloads11mo agoHugging Face10krishagarwal /evaluation1tabularn<1K0 likes85 downloads11mo agoHugging Face11krishagarwal /evaluation2tabularn<1K0 likes85 downloads11mo agoHugging Face12rasinmuhammed /ecommerce-analytics-sql-evaluation Ecommerce Analytics SQL Evaluation (declared GMV, verified answer key) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/ecommerce-analytics-sql-evaluation.tabulartable-question-answering10K<n<100K0 likes80 downloads2mo agoHugging Face13nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes78 downloads20d agoHugging Face14deutsche-telekom /NLU-Evaluation-Data-en-de NLU Evaluation Data - English and German A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction. This dataset is collected and annotated for evaluating NLU services and platforms. The detailed paper on this dataset can be found at arXiv.org: Benchmarking Natural Language Understanding Services for building Conversational Agents The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.tabulartext-classification10K<n<100K2 likes66 downloads3y agoHugging Face15CQA-pharma /cqa-ai-technical-response-evaluation CQA AI Technical Response Evaluation Dataset Overview This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses. The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.tabulartext-classificationn<1K0 likes63 downloads15d agoHugging Face16fai-adh /fon-code-switching-evaluation French-Fon Code-Switching Evaluation Benchmark Overview This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios. The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon. The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.tabularquestion-answeringn<1K0 likes62 downloads2d agoHugging Face17RicardoRei /wmt-sqm-human-evaluation Dataset Summary In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA). The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.tabular100K<n<1M1 likes59 downloads4y agoHugging Face18Cross-Mergeability /extrinsic-evaluations Extrinsic evaluations — the union view One tidy long-format table of every extrinsic (downstream, task-level) evaluation produced across the 2026-08-26 mergeability workstreams, so that a single file answers "how did model X score on benchmark Y" regardless of which experiment produced it. The per-experiment datasets remain the authoritative record of their own methods, figures and caveats. This is the union view, not a replacement, and it deliberately carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.tabular1K<n<10K0 likes53 downloads1mo agoHugging Face19rasinmuhammed /saas-finance-sql-evaluation SaaS Finance SQL Evaluation (MRR waterfalls that reconcile exactly) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/saas-finance-sql-evaluation.tabulartable-question-answering1K<n<10K0 likes46 downloads2mo agoHugging Face20ucberkeley-dlab /normative_evaluation_llms_everyday_dilemmastabular10K<n<100K2 likes45 downloads1y agoHugging Face21vicharanashala-org /continuous-evaluation-in-programming-course-logs Continuous Evaluation in Programming Course - Faculty and Student Logs Two tables from a first-year programming course run at National Institute in India, July to December 2023. Students worked through a ladder of 60 programming tasks (O01 to O60), marked each task complete themselves on a course dashboard, and a teaching assistant then checked every claimed task in a short viva. Files self_reported_claims.csv — 203 rows, one per student. column meaning… See the full description on the dataset page: https://huggingface.co/datasets/vicharanashala-org/continuous-evaluation-in-programming-course-logs.tabularn<1K2 likes42 downloads7d agoHugging Face22rasinmuhammed /edtech-sql-evaluation EdTech SQL Evaluation (declared learning curve, verified answer key) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/edtech-sql-evaluation.tabulartable-question-answering10K<n<100K0 likes40 downloads2mo agoHugging Face23sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes31 downloads1y agoHugging Face24smallstepai /Misal-Evaluation-v0.1tabularn<1K0 likes28 downloads2y agoHugging Face25AGundawar /chess_position_evaluationstabular10M<n<100M0 likes27 downloads2y agoHugging Face26patchmedia-org /remote-ai-evaluation-training-market-snapshot Dataset Description This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities. The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory. Reporting date: 2026-08-22 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.tabularn<1K0 likes27 downloads2mo agoHugging Face27g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes25 downloads2mo agoHugging Face28ismielabir /Quantum_Gate_Performance_Evaluation 🧪 Quantum Gate Performance Dataset 📘 Title: Comprehensive Quantum Gate Performance Analysis: A Comparative Study of Noise and No-Noise Effects 📂 Dataset Description: This repository contains benchmarking results for 13 quantum gates (e.g., H, CNOT, Toffoli) tested under noisy and noise-free conditions, based on 1000 simulation runs per gate configuration. Total 26000 rows and 13 columns. 📊 Features include: Gate Type Execution Time Error Rate Fidelity… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Quantum_Gate_Performance_Evaluation.tabular10K<n<100K0 likes21 downloads1y agoHugging Face29jason1966 /algozee_mas-feature-evaluation-dataset MAS Feature Evaluation Dataset A Structured Dataset for Multi-Agent Capability Analysis Dataset Info Source: Kaggle Original Size: 0.01 MB Kaggle Downloads: 14 Files: 1 Files mas_dataset.csv Mirrored from Kaggle tabular1K<n<10K0 likes21 downloads6mo agoHugging Face30Equall /perplexity_evaluation SaulLM-7B: Pioneering the first Legal Large Language Model Perplexity Analysis This dataset presents the data used in the paper "SaulLM-7B: Pioneering the first Legal Large Language Model" in "6.3 Perplexity Analysis" section. The dataset contains the perplexity scores of SaulLM-7B, Llama2-7B and Mistral-7B across a corpora of recent text. Cleaning We proceeded to standardize the data by removing any special characters using unicodedata normalization. We also… See the full description on the dataset page: https://huggingface.co/datasets/Equall/perplexity_evaluation.tabular1K<n<10K3 likes19 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.