Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uclanlp /OpenVLHarness-Evaluation-Datasets OpenVLHarness evaluation datasets Processed evaluation splits used by OpenVLHarness (project page). Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.imagevisual-question-answering10K<n<100K1 likes1.6k downloads1d agoHugging Face02CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes660 downloads5mo agoHugging Face03TechWolf /JobBERT-evaluation-dataset JobBERT evaluation dataset 💾 This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation. The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board. Access the JobBERT paper here: https://arxiv.org/abs/2109.09605 BibTeX Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.texttext-classification10K<n<100K0 likes560 downloads1y agoHugging Face04RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes420 downloads4y agoHugging Face05allganize /RAG-Evaluation-Dataset-KO Allganize RAG Leaderboard Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다. RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다. RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다. Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.textn<1K123 likes378 downloads2y agoHugging Face06allganize /RAG-Evaluation-Dataset-JA Allganize RAG Leaderboard とは Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.textn<1K42 likes285 downloads2y agoHugging Face07egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face08RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes217 downloads4y agoHugging Face09felixleungsc /paperswithcode-data-evaluation-tables Process data from paperswithcode See https://huggingface.co/datasets/pwc-archive/files/tree/main. Download and unzip evaluation tables: curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz" gunzip jul-28-evaluation-tables.json.gz Install jq. See https://jqlang.org/. If on Debian/Ubuntu, install with sudo apt-get install jq. Example jq to extract: jq -r ' def process(parent): .task as $current_task | (if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.text100K<n<1M1 likes212 downloads1y agoHugging Face10OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes206 downloads3y agoHugging Face11recruit-jp /japanese-image-classification-evaluation-dataset recruit-jp/japanese-image-classification-evaluation-dataset Overview Developed by: Recruit Co., Ltd. Dataset type: Image Classification Language(s): Japanese LICENSE: CC-BY-4.0 More details are described in our tech blog post. 日本語CLIP学習済みモデルとその評価用データセットの公開 Dataset Details This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks. jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.imageimage-classification1K<n<10K10 likes155 downloads3y agoHugging Face12JosephAA /toporisk-evaluation-artifacts TopoRisk Evaluation Artifacts English | 简体中文 TopoRisk studies how a limited verification budget should be allocated across an agentic workflow graph. Instead of ranking steps only by their local failure probability, the scheduler estimates how an error can reach terminal outputs and recomputes marginal value after every selected audit. This repository is an artifact-first research preview. It contains code, aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.tabularn<1K1 likes155 downloads2d agoHugging Face13chillies /IELTS-writing-task-2-evaluationtext10K<n<100K39 likes149 downloads3y agoHugging Face14Momina-Muzafar /urdu-english-llm-evaluation Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories. Motivation It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.textquestion-answeringn<1K0 likes148 downloads8d agoHugging Face15CrimsonRobot /crimson-ma2-evaluation Crimson MolmoAct2: preliminary real-robot evaluation Right-arm bottle pick-and-place rollouts recorded on 2026-09-24 (Asia/Bangkok). This release contains three checkpoints, 31 logged attempts, videos from three cameras, selected recorded telemetry, and joint trajectories. It is an evaluation evidence release, not a training dataset or a completed six-checkpoint benchmark. Model source: Kavin60606/crimson-ma2-ckpts. The policy task text was pick and place the bottle. Each… See the full description on the dataset page: https://huggingface.co/datasets/CrimsonRobot/crimson-ma2-evaluation.tabularroboticsn<1K0 likes92 downloads16d agoHugging Face16krishagarwal /evaluation5tabularn<1K0 likes86 downloads11mo agoHugging Face17krishagarwal /evaluation1tabularn<1K0 likes85 downloads11mo agoHugging Face18krishagarwal /evaluation2tabularn<1K0 likes85 downloads11mo agoHugging Face19rasinmuhammed /ecommerce-analytics-sql-evaluation Ecommerce Analytics SQL Evaluation (declared GMV, verified answer key) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/ecommerce-analytics-sql-evaluation.tabulartable-question-answering10K<n<100K0 likes80 downloads2mo agoHugging Face20nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes78 downloads20d agoHugging Face21WueNLP /mHallucination_Evaluation Multilingual Hallucination Evaluation in the wild The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset): Dataset Details The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.texttext-generation10K<n<100K0 likes77 downloads2y agoHugging Face22deutsche-telekom /NLU-Evaluation-Data-en-de NLU Evaluation Data - English and German A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction. This dataset is collected and annotated for evaluating NLU services and platforms. The detailed paper on this dataset can be found at arXiv.org: Benchmarking Natural Language Understanding Services for building Conversational Agents The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.tabulartext-classification10K<n<100K2 likes66 downloads3y agoHugging Face23duaa-amer /ai-blind-spot-evaluation-dataset Evaluating Diagnostic Recognition Blind Spots in Small Models: A Cross-Model Study of Endometriosis Note: this is the dataset for the Fatima Fellowship Application Fall 2026 technical challenge. Answer to Question 1 and Question 3 Healthcare practitioners, when faced with atypical manifestations of diseases, often resort to hypothesis generation and revision if classic pattern recognition doesn’t suffice. This raises the question that sits at the heart of this evaluation: How… See the full description on the dataset page: https://huggingface.co/datasets/duaa-amer/ai-blind-spot-evaluation-dataset.textn<1K0 likes65 downloads15h agoHugging Face24awaaz-se-alfaaz /YouTube-Evaluation-Set Awaaz se Alfaaz — YouTube Evaluation Set This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.textautomatic-speech-recognitionn<1K0 likes63 downloads2mo agoHugging Face25nosuke113 /libero-plus-evaluation LIBERO-plus Evaluation Dataset (180 held-out tasks) This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark. It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models: Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise. View — Camera… See the full description on the dataset page: https://huggingface.co/datasets/nosuke113/libero-plus-evaluation.textn<1K0 likes63 downloads16d agoHugging Face26CQA-pharma /cqa-ai-technical-response-evaluation CQA AI Technical Response Evaluation Dataset Overview This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses. The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.tabulartext-classificationn<1K0 likes63 downloads15d agoHugging Face27fai-adh /fon-code-switching-evaluation French-Fon Code-Switching Evaluation Benchmark Overview This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios. The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon. The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.tabularquestion-answeringn<1K0 likes62 downloads2d agoHugging Face28plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes61 downloads1mo agoHugging Face29RicardoRei /wmt-sqm-human-evaluation Dataset Summary In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA). The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.tabular100K<n<1M1 likes59 downloads4y agoHugging Face30nwhite-systems /responsible-agent-workflow-evaluation Responsible Agent Workflow Evaluation Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating whether an AI agent respects safety, permission and accountability boundaries in operational settings. Thirteen categories contain ten scenarios each. Every record includes an intentionally unsafe request, contextual facts, expected safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria and reviewer guidance. This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.textn<1K1 likes53 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.