Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.5k downloads1y agoHugging Face02zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads5mo agoHugging Face03tsinghua-sigs-robot-lab /VeriLoop-E2-Evaluation-Evidence VeriLoop E2 Evaluation Evidence Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks. This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results. The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.text-generation0 likes1.8k downloads13d agoHugging Face04JesseLiu /patient-evaluations Patient Evaluations Dataset This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data. Dataset Description The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions. Dataset Structure The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.texttext-generationn<1K0 likes675 downloads8mo agoHugging Face05mtec-TUB /GPT-4o-evaluation-biases A database to support the evaluation of gender biases in GPT-4o output The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025). Introduction This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.question-answering10K<n<100K0 likes500 downloads2y agoHugging Face06dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes351 downloads12d agoHugging Face07YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes324 downloads5mo agoHugging Face08aisingapore /Cultural-Evaluation-Kalahigated Kalahi Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM. Supported Tasks and Leaderboards Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.textmultiple-choicen<1K0 likes298 downloads9mo agoHugging Face09TsinghuaC3I /ZEDA-Evaluation ZEDA Dataset This repository contains the training data for ZEDA (Zero-Expert Self-Distillation Adaptation), a framework introduced in the paper Post-Trained MoE Can Skip Half Experts via Self-Distillation. ZEDA is a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones by injecting zero experts and using self-distillation. Paper: Post-Trained MoE Can Skip Half Experts via Self-Distillation GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/ZEDA-Evaluation.texttext-generation0 likes224 downloads5mo agoHugging Face10MultiverseComputingCAI /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.texttext-generation1K<n<10K5 likes223 downloads10mo agoHugging Face11sreearravind /AI-Research-Evaluation-Repository-STEM AI-STEM-Research-Eval-Dataset Overview This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations. It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content. The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.text-generationn<1K1 likes199 downloads4mo agoHugging Face12Dynamicresponselabs /JASON-High-Stakes-AI-Evaluation-Samples J.A.S.O.N. Evaluation Sample Previews V01-V29 Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts. The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.texttext-generationn<1K0 likes163 downloads13d agoHugging Face13compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes156 downloads5mo agoHugging Face14Momina-Muzafar /urdu-english-llm-evaluation Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories. Motivation It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.textquestion-answeringn<1K0 likes148 downloads8d agoHugging Face15LaurelWings /rcga-evaluation-data RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.tabulartext-generation10K<n<100K0 likes134 downloads11d agoHugging Face16dougdotcon /douvras-ptbr-enterprise-ai-evaluation Douvras PT-BR Enterprise AI Evaluation Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada: responder somente a partir de um documento fornecido; reconhecer quando a informação não está disponível; resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.textquestion-answering0 likes133 downloads28d agoHugging Face17arjhinety /OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation OpenGrad — Qwen3.5-2B, M0 SFT on corpus v1: evaluation record This repository holds the evaluation evidence for one OpenGrad experiment, qwen35_2b_m0_sft_full_v3: a full-parameter supervised fine-tuning run of Qwen/Qwen3.5-2B on the published OpenGrad ToolPolicy Canonical v1 corpus. There are no model weights here, and none exist. Every checkpoint this run produced was deleted from local storage before it was uploaded, and none of them can be recovered. This repository is what… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation.text-generation0 likes127 downloads17d agoHugging Face18egeboy35 /qwen3.5-4b-tr-en-evaluation#Qwen3.5-4B: Turkish and English Responses #Model Selection / Main Findings If we look at the Turkish performance of Qwen3.5-4B in multilingual settings, considering the values on Hugging Face, it has approximately 80% accuracy on 120 questions in TR-MMLU-like knowledge questions, but 33.3% in the Turkish culture and idioms section; it has 47.5% accuracy in TR-HellaSwag-like completion tasks. When we examine it overall, it achieves 65% accuracy with a limited sample (360 questions). However… See the full description on the dataset page: https://huggingface.co/datasets/egeboy35/qwen3.5-4b-tr-en-evaluation.text-generation0 likes105 downloads2d agoHugging Face19beatsprom /stateless-mcp-agent-evaluation-suite-2026 ⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1 ⚡ Overview & Industry Problem As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.tabulartext-generation1K<n<10K0 likes95 downloads18d agoHugging Face20Lots-of-LoRAs /task1338_peixian_equity_evaluation_corpus_sentiment_classifier Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.texttext-generation1K<n<10K0 likes91 downloads2y agoHugging Face21SHPDRG /medical-imaging-model-evaluation-benchmark 医学影像多模态模型评测集(精选示例版) 这是一个面向医学多模态大模型的高质量影像评测集,专门测试模型能否把“看见影像”进一步转化为可解释、可复核、符合临床语境的判断与表达。数据将医学影像与患者描述、病史摘要、检查信息或结构化临床资料配对,覆盖从影像分类、报告生成,到鉴别诊断、治疗方案和胸片质量控制的完整评测链路。 本次公开版本从 2026-07-22 质检通过产物中整理而来,按每个子集最多 50 题进行分层抽样;题量不足 50 的影像质量控制子集完整保留。因此,公开版本包含 5 个任务子集、213 题和 455 个配套影像文件,适合作为医学视觉语言模型的快速对比集、回归测试集和研究教学样例。 数据集亮点 多模态对齐:每条样例同时提供影像和结构化的 question、answer、explanation,支持检查视觉理解、临床语义整合与解释质量。 任务覆盖完整:从“影像是什么”到“如何描述、如何鉴别、如何处置”,并加入真实影像工作流中的胸片质量控制任务。 影像类型丰富:覆盖 X… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-imaging-model-evaluation-benchmark.imageimage-classificationn<1K0 likes86 downloads3d agoHugging Face22WueNLP /mHallucination_Evaluation Multilingual Hallucination Evaluation in the wild The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset): Dataset Details The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.texttext-generation10K<n<100K0 likes77 downloads2y agoHugging Face23openbmb /RLPR-Evaluation Dataset Card for RLPR-Evaluation GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.textquestion-answeringn<1K3 likes77 downloads1y agoHugging Face24Lux0926 /ASPRM-BON-Evaluation-Dataset-CodeThis repository contains the data from the paper AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence. Source code: https://github.com/Lux0926 text-generation0 likes67 downloads2y agoHugging Face25CQA-pharma /cqa-ai-technical-response-evaluation CQA AI Technical Response Evaluation Dataset Overview This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses. The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.tabulartext-classificationn<1K0 likes63 downloads15d agoHugging Face26DBbun /synthetic-llm-evaluation-traces EvaluLLM-Inspired Synthetic Evaluation Traces (DBbun) View source code on GitHub Watch on Youtube: Evaluating AI with AI Dataset Summary This dataset contains fully synthetic evaluation traces for NLG / LLM-style output comparison, inspired by the evaluation workflow described in EvaluLLM: LLM Assisted Evaluation of Generative Outputs (IUI Companion 2024). The dataset is produced by a configurable, offline simulator and includes: synthetic tasks (prompts) synthetic… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/synthetic-llm-evaluation-traces.text-generation0 likes59 downloads9mo agoHugging Face27Astride1 /frontier-model-blindspot-evaluation Technical Challenge: Blind Spots of Frontier Models in Operational Physics and Distributed Microgrids Candidate: Astride Melvin Fokam NinyimModel Evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.54B Parameters)Artifacts: Evaluation notebook (eval_notebook.ipynb) and empirical outputs (results_log.json). Part 1: Blind Spot & Capability Gap Identification Drawing directly from my lived experience and research background in Sub-Saharan Africa and cyber-physical energy… See the full description on the dataset page: https://huggingface.co/datasets/Astride1/frontier-model-blindspot-evaluation.text-generationn<1K0 likes56 downloads19d agoHugging Face28AnjanSB /NQ-RAG-DPO-Evaluation Dataset Card Dataset Summary This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO). The system is organized into three interconnected pipelines: 1️. RAG Pipeline The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark. For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.texttext-generation1K<n<10K1 likes53 downloads8mo agoHugging Face29Inkwell-Software /screenplay-revision-evaluation Screenplay Revision Evaluation Cases 24 original screenwriting revision tasks. Each gives a short scene and a constraint — cut a page to its beat, plant a prop, hold an answer back, fix a continuity slip — then pairs it with mechanical checks (a word ceiling, a line that must survive) and separate human-review questions. It tests whether a tool, or a person, can make a tightly-constrained edit while keeping the scene intact. Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.texttext-generationn<1K0 likes53 downloads16d agoHugging Face30Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes52 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.