Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.5k downloads1y agoHugging Face02mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K39 likes2.8k downloads3y agoHugging Face03Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 416,442,401 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 1,003,347,246 rows. This dataset is updated monthly, and was last updated on October 7th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular1B<n<10B33 likes2.6k downloads3d agoHugging Face04uclanlp /OpenVLHarness-Evaluation-Datasets OpenVLHarness evaluation datasets Processed evaluation splits used by OpenVLHarness (project page). Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.imagevisual-question-answering10K<n<100K1 likes1.6k downloads18h agoHugging Face05ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes897 downloads2y agoHugging Face06CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes660 downloads5mo agoHugging Face07RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes420 downloads4y agoHugging Face08ymoslem /wmt-da-human-evaluation-long-context Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.tabular1M<n<10M9 likes352 downloads2y agoHugging Face09dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes351 downloads12d agoHugging Face10YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes324 downloads5mo agoHugging Face11demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes310 downloads2y agoHugging Face12josancamon /kg-gen-MINE-evaluation-datasettabularn<1K6 likes277 downloads1y agoHugging Face13Geraldine /humatheque-vlm-evaluation-results Humathèque — VLM metadata-extraction evaluation Structured metadata extraction from French thesis/dissertation title pages, scored against a human-validated visible-only ground truth (catalogue role/jury data not on the page was removed). Leaderboard model overall surface stable role valid_json gemma4-26b-a4b 0.947 0.8 0.942 0.961 1 mistral-small-32-24b 0.939 0.728 0.925 0.98 1 bonsai2-27b 0.939 0.687 0.934 0.95 1 lift 0.931 0.51 0.92 0.96 1… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/humatheque-vlm-evaluation-results.tabular10K<n<100K0 likes264 downloads3d agoHugging Face14togethercomputer /CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriestabularn<1K13 likes251 downloads8mo agoHugging Face15egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face16RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes217 downloads4y agoHugging Face17MERA-evaluation /SWE-MERA SWE-MERA Continuously updated SWE-MERA dataset SWE-MERA splits: dev: for testing (10 samples) lite: presented at the leaderboard here (750 samples) full: continuously updated to collect more data (2738 samples) Load dataset from datasets import load_dataset ds = load_dataset("MERA-evaluation/SWE-MERA", split='dev') Evaluation Description The main tool to validate tasks is repotest (available at PyPI or GitHub) data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/SWE-MERA.tabularother1K<n<10K11 likes213 downloads9mo agoHugging Face18OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes206 downloads3y agoHugging Face19compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes156 downloads5mo agoHugging Face20JosephAA /toporisk-evaluation-artifacts TopoRisk Evaluation Artifacts English | 简体中文 TopoRisk studies how a limited verification budget should be allocated across an agentic workflow graph. Instead of ranking steps only by their local failure probability, the scheduler estimates how an error can reach terminal outputs and recomputes marginal value after every selected audit. This repository is an artifact-first research preview. It contains code, aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.tabularn<1K1 likes155 downloads2d agoHugging Face21bdatm-project /evaluation-results-task1Model-comparison table for this task: one row per evaluated model, written by push_results_table in src/eval/utilities.py. Per-sample predictions are in per_sample/. tabularn<1K0 likes153 downloads8d agoHugging Face22Scicom-intl /Evaluation-Multilingual-VC Evaluation-Multilingual-VC We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon, Filter languages that support by Whisper Large V3 to evaluate WER automatically, Only take test set, sort by up votes. Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows. Only build first 500 rows for each language Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.audio10K<n<100K0 likes151 downloads6mo agoHugging Face23YanZhanPKU /Byte-Authority-Evaluation Byte Authority · Evaluation Records 📄 Paper (arXiv:2609.35932) &nbsp; • &nbsp; 💻 Code &nbsp; • &nbsp; 🤗 Collection This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.tabular100K<n<1M1 likes151 downloads10d agoHugging Face24AITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K4 likes139 downloads3mo agoHugging Face25datalama /RAG-Evaluation-Dataset-KO Dataset Card for Reconstructed RAG Evaluation Dataset (KO) Dataset Summary 본 데이터셋은 allganize/RAG-Evaluation-Dataset-KO를 기반으로 PDF 파일을 포함하도록 재구성한 한국어 평가 데이터셋입니다. 원본 데이터셋에서는 PDF 파일의 경로만 제공되어 수동으로 파일을 다운로드해야 하는 불편함이 있었고, 일부 PDF 파일의 경로가 유효하지 않은 문제를 보완하기 위해 PDF 파일을 포함한 데이터셋을 재구성하였습니다. Supported Tasks and Leaderboards RAG Evaluation: 본 데이터는 한국어 RAG 파이프라인에 대한 E2E Evaluation이 가능합니다. Languages The dataset is in Korean (ko). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/datalama/RAG-Evaluation-Dataset-KO.tabularother1K<n<10K0 likes137 downloads2y agoHugging Face26bdatm-project /evaluation-results-task2Model-comparison table for this task: one row per evaluated model, written by push_results_table in src/eval/utilities.py. Per-sample predictions are in per_sample/. tabularn<1K0 likes137 downloads8d agoHugging Face27LaurelWings /rcga-evaluation-data RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.tabulartext-generation10K<n<100K0 likes134 downloads10d agoHugging Face28ProlificAI /humaine-evaluation-dataset HUMAINE: Human-AI Interaction Evaluation Dataset Dataset Description Dataset Summary The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases. The dataset consists of two main components: Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.tabularquestion-answering100K<n<1M6 likes129 downloads5mo agoHugging Face29Cyber-security-final-project /Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition Injected PDFs - Model Evaluation This repository holds the model evaluation stage of a project on detecting harmless-but-real attack payloads injected into PDF files, together with the artefacts it produced for the application. Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs, and the two winners are exported for the app to load. Question Candidates Winner Part A Which files look like this one? 3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.tabulartext-classification1K<n<10K0 likes117 downloads2mo agoHugging Face30bdatm-project /evaluation-results-task3Model-comparison table for this task: one row per evaluated model, written by push_results_table in src/eval/utilities.py. Per-sample predictions are in per_sample/. tabularn<1K0 likes106 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.