Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K39 likes2.8k downloads3y agoHugging Face02dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes351 downloads12d agoHugging Face03demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes310 downloads2y agoHugging Face04compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes156 downloads5mo agoHugging Face05YanZhanPKU /Byte-Authority-Evaluation Byte Authority · Evaluation Records 📄 Paper (arXiv:2609.35932) &nbsp; • &nbsp; 💻 Code &nbsp; • &nbsp; 🤗 Collection This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.tabular100K<n<1M1 likes151 downloads11d agoHugging Face06AITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K4 likes139 downloads3mo agoHugging Face07LaurelWings /rcga-evaluation-data RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.tabulartext-generation10K<n<100K0 likes134 downloads11d agoHugging Face08Usman391 /WildChat-Evaluation-Awareness-Taxonomy What do models think WildChat is testing? This is the dataset behind the exploratory blog “What do models think WildChat is testing?”, written by Codex (GPT-6-Astra-XHigh), guided by Usman Anwar. It contains 766 VEA-positive responses from Qwen, Inkling, and Kimi K2.6, with full initial user prompts, saved reasoning, and GPT-5.6-Luna taxonomy labels. VEA means verbalized evaluation awareness: the model says or considers that it is being tested. Methodology We used… See the full description on the dataset page: https://huggingface.co/datasets/Usman391/WildChat-Evaluation-Awareness-Taxonomy.tabulartext-classification10K<n<100K1 likes99 downloads9d agoHugging Face09ZurichNLP /romansh-mt-evaluation Dataset Description This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists. The evaluation covers three quality dimensions: Document accuracy, in which annotators assessed the adequacy of complete document translations. Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.tabular1K<n<10K0 likes52 downloads3mo agoHugging Face10wayne-redemption /Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset 📌 Dataset Contents Each sample includes: category: The evaluation domain prompt: The question given to the LLM temperature: Environmental temperature input humidity: Environmental humidity input context: A scenario label (e.g., cool_humid, hot_dry, average_day) reference: Expert-crafted expected output All data is provided in a single JSON file. 🧪 Intended Use This dataset supports research on: LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.tabulartext-classificationn<1K0 likes45 downloads11mo agoHugging Face11compass-group-tue /sdf_evaluation_traits_15M Models That Know How Evaluations Are Designed Score Safer This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.tabulartext-generation10K<n<100K0 likes39 downloads2mo agoHugging Face12speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes35 downloads2y agoHugging Face13sentinelseed /sentinel-evaluations Sentinel Evaluations Evaluation results for multiple alignment seeds across various AI safety benchmarks. Overview This dataset contains: Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI) Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more Quick Start from datasets import load_dataset # Load seeds seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train") # Load results results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.tabulartext-classificationn<1K0 likes34 downloads10mo agoHugging Face14SeanWang0027 /data_for_evaluationtabular1K<n<10K0 likes25 downloads6mo agoHugging Face15Debbyjaye001 /adaption-marketing-fit-evaluations This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-marketing_fit_evaluations This dataset contains prompt-completion pairs where an AI evaluates marketing communications for specific regional markets, focusing on trust, cultural context, and decision friction. Each completion provides a market fit score, strategic diagnosis, and a rewritten copy optimized for the target audience and platform. The entries cover diverse sectors and… See the full description on the dataset page: https://huggingface.co/datasets/Debbyjaye001/adaption-marketing-fit-evaluations.tabular10K<n<100K0 likes16 downloads3mo agoHugging Face16parsa-mhmdi /RadLLamaThinking_Stage2_Final_Evaluationtabularn<1K0 likes11 downloads1y agoHugging Face17trikwi /chsa-p14-sealed-evaluation chsa-p14-sealed-evaluation Jeu d'évaluation scellé — ouvert une seule fois, à la toute fin. Ne jamais utiliser pour entraîner ni pour régler. Ce dépôt ne doit jamais servir à entraîner ni à régler un modèle. Il est ouvert une seule fois, à la fin, sur tous les états du modèle à la fois, avec le même gabarit et la même configuration de génération. Les décisions de promotion se prennent avant, sur le jeu de validation. Projet CHSA P14 — preuve de concept d'un agent de triage… See the full description on the dataset page: https://huggingface.co/datasets/trikwi/chsa-p14-sealed-evaluation.tabular1K<n<10K0 likes10 downloads2mo agoHugging Face18wvnvwn /Downstream-Evaluation-SIA Downstream-Evaluation-SIA Fixed downstream evaluation subsets for SIA experiments. Configs mmlu: cais/mmlu, 57 subject subsets, 50 test examples per subject sampled with seed 42. Total: 2850. spider: xlangai/spider, prepared seed 42 test split from the validation split. Total: 517. pubmedqa: qiaojin/PubMedQA, pqa_artificial, prepared seed 42 test split. Total: 500. The original JSON snapshots with metadata are also stored under snapshots/. tabular1K<n<10K0 likes9 downloads5mo agoHugging Face19Noothi /telugu-indicf5-evaluationaudion<1K0 likes7 downloads2mo agoHugging Face20parsa-mhmdi /RadLLamaThinking_Stage1_Final_Evaluationtabularn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.