Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads5mo agoHugging Face02mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K39 likes2.8k downloads3y agoHugging Face03MERA-evaluation /MERA MERA v.1.2.0 (Multimodal Evaluation for Russian-language Architectures) Summary MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language. The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.text10K<n<100K11 likes2.4k downloads2d agoHugging Face04FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes626 downloads2y agoHugging Face05PKU-Alignment /BeaverTails-Evaluation Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.texttext-classificationn<1K15 likes545 downloads3y agoHugging Face06WenyiWU0111 /webvoyager_evaluation_datatextn<1K0 likes536 downloads1y agoHugging Face07dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes351 downloads12d agoHugging Face08demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes310 downloads2y agoHugging Face09ZeeshanSaud /CodeTruthAgent-V3-Module1-Evaluation Source Code https://github.com/Zeeshan78699/CodeTruthAgent — tag v3.0.0-module1 CodeTruth Agent V3 — Module 1 Evaluation Validation results for Module 1: Repository Cognition Engine — a deterministic, rule-based engine that scans a software repository and determines its application type, primary framework, technology stack, and file inventory. What's in this dataset FULL_DOMAIN_SUMMARY.md — summary table of all 69 validated repositories… See the full description on the dataset page: https://huggingface.co/datasets/ZeeshanSaud/CodeTruthAgent-V3-Module1-Evaluation.textn<1K0 likes183 downloads4mo agoHugging Face10phionyx /airep-embedded-evaluation-profile AIREP Embedded Evaluation Profile v0.1 This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical specification history lives in the AIREP GitHub repository. Byte identity between this mirror and its source commit is a distribution-integrity property, not independent scientific verification. Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.textn<1K0 likes165 downloads23d agoHugging Face11Dynamicresponselabs /JASON-High-Stakes-AI-Evaluation-Samples J.A.S.O.N. Evaluation Sample Previews V01-V29 Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts. The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.texttext-generationn<1K0 likes163 downloads13d agoHugging Face12compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes156 downloads5mo agoHugging Face13anikethh /Comparative-Idea-Evaluation 🔬 Comparative Idea Evaluation Dataset accompanying our Findings of ACL 2026 paper: Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation. Srujan P Mule · Aniketh Garikaparthi · Manasi Patwardhan 📄 ACL Anthology · arXiv Research overview from Figure 1 of the paper. This release contains the comparison datasets; reasoning-training variants illustrated in the figure are not included. 🧠 What is this dataset for? Given a research… See the full description on the dataset page: https://huggingface.co/datasets/anikethh/Comparative-Idea-Evaluation.texttext-classification10K<n<100K0 likes154 downloads24d agoHugging Face14YanZhanPKU /Byte-Authority-Evaluation Byte Authority · Evaluation Records 📄 Paper (arXiv:2609.35932) &nbsp; • &nbsp; 💻 Code &nbsp; • &nbsp; 🤗 Collection This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.tabular100K<n<1M1 likes151 downloads10d agoHugging Face15AITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K4 likes139 downloads3mo agoHugging Face16orhunc /Bias-Evaluation-TurkishTranslation of bias evaluation framework of May et al. (2019) from this repository and this paper into Turkish. There is a total of 37 tests including tests addressing gender-bias as well as tests designed to evaluate the ethnic bias toward Kurdish people in Türkiye context. Abstract of the paper: While the growing size of pre-trained language models has led to large improvements in a variety of natural language processing tasks, the success of these models comes with a price: They are trained… See the full description on the dataset page: https://huggingface.co/datasets/orhunc/Bias-Evaluation-Turkish.textn<1K1 likes136 downloads4y agoHugging Face17LaurelWings /rcga-evaluation-data RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.tabulartext-generation10K<n<100K0 likes134 downloads10d agoHugging Face18aliahmar /udhaar-evaluationtextn<1K0 likes124 downloads8d agoHugging Face19ltg /normistral-11b-thinking-evaluationtext10K<n<100K1 likes107 downloads10mo agoHugging Face20s-emanuilov /rivers-evaluation-results Rivers Evaluation Results - Comprehensive LLM Benchmarking All results from the paper's five experimental conditions: baseline LLMs, fine-tuned models, RAG, and Graph-RAG with Licensing Oracle. This repository contains baseline evaluations for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and Gemma 3-4B, along with fine-tuning results for both factual recall and abstention behavior. It also includes outputs from the embedding-based RAG system and the Graph-RAG with Licensing Oracle… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-evaluation-results.text10K<n<100K0 likes101 downloads11mo agoHugging Face21CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes100 downloads6mo agoHugging Face22Usman391 /WildChat-Evaluation-Awareness-Taxonomy What do models think WildChat is testing? This is the dataset behind the exploratory blog “What do models think WildChat is testing?”, written by Codex (GPT-6-Astra-XHigh), guided by Usman Anwar. It contains 766 VEA-positive responses from Qwen, Inkling, and Kimi K2.6, with full initial user prompts, saved reasoning, and GPT-5.6-Luna taxonomy labels. VEA means verbalized evaluation awareness: the model says or considers that it is being tested. Methodology We used… See the full description on the dataset page: https://huggingface.co/datasets/Usman391/WildChat-Evaluation-Awareness-Taxonomy.tabulartext-classification10K<n<100K1 likes99 downloads8d agoHugging Face23gauravshrm211 /VC-startup-evaluation-for-investmentThis data set includes the completion pairs for evaluating startups before investing in them. This data set iincludes completion examples for Chain of Thought reasoning to perform financial calculations. This data set includes completion examples for evaluating risk profile, growth propspects, cost, ratios, market size, asset, liability, debt, equity and other ratios. This data set includes comparison of different startups. textn<1K13 likes95 downloads3y agoHugging Face24bingbangboom /stockfish-evaluation-SAN Dataset Card for the Stockfish Evaluations A dataset of chess positions evaluated with various flavours of Stockfish running within user browsers. Produced by, and for, the Lichess analysis board. Evaluations are formatted as JSON; one position per line. The schema of a position looks like this: { "fen": "8/8/2B2k2/p4p2/5P1p/Pb6/1P3KP1/8 w - -", "depth": 42, "evaluation": 5.64, "best_move": "Kg1", "best_line": "Kg1 Ke6 Kh2 Kd6 Be8 Kc5 Kh3 Kd6 Bb5 Ke7" } fen: string, the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/stockfish-evaluation-SAN.text10M<n<100M3 likes82 downloads2y agoHugging Face25ValidatorMasterHeartcodeProtocol /heartcode-evaluation-fixtures Heartcode Experimental Action Evidence v0.1 Eight original synthetic fixtures for an experimental action-submission boundary. GitHub is canonical: release directory. Hugging Face is a distribution copy of these exact files. What this measures A host-owned test grant allows exactly one bound RESET proposal. Missing, expired, revoked, stale, mismatched or unverifiable authority blocks submission to a fake environment. Expected records include deterministic decisions… See the full description on the dataset page: https://huggingface.co/datasets/ValidatorMasterHeartcodeProtocol/heartcode-evaluation-fixtures.textothern<1K1 likes81 downloads25d agoHugging Face26minhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes76 downloads4mo agoHugging Face27Reza-Telus /certainty-robustness-llm-evaluation Certainty Robustness Benchmark This repository accompanies the paper: Certainty robustness: Evaluating LLM stability under self-challenging promptsMohammadreza Saadat, Steve NemzerarXiv:2603.03330, 2026https://arxiv.org/abs/2603.03330 Overview The Certainty Robustness Benchmark evaluates how large language models (LLMs) behave when their initial answers are challenged by follow-up prompts such as: “Are you sure?” “You are wrong!” confidence elicitation prompts Rather… See the full description on the dataset page: https://huggingface.co/datasets/Reza-Telus/certainty-robustness-llm-evaluation.documentn<1K0 likes71 downloads7mo agoHugging Face28RKB109 /rag-evaluation-lab-20260928-dataset RAG Evaluation Lab Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for RAG systems often ship without a stable regression set or failure taxonomy. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant: generation pattern synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260928-dataset.texttext-classificationn<1K0 likes68 downloads12d agoHugging Face29CK0607 /komorebi-painter-evaluationstextn<1K0 likes66 downloads13d agoHugging Face30RKB109 /rag-evaluation-lab-20260918-dataset RAG Evaluation Lab Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for RAG systems often ship without a stable regression set or failure taxonomy. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant: generation pattern synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260918-dataset.texttext-classificationn<1K0 likes65 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.