Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anon-cmevs-2026 /cmevs-erp-eval CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution. v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.imagedepth-estimationn<1K10 likes51k downloads4mo agoHugging Face02MVU-Eval-Team /MVU-Eval-Data MVU-Eval Dataset Paper | Code | Project Page Dataset Description The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.tabularvideo-text-to-text1K<n<10K2 likes3.8k downloads11mo agoHugging Face03evaluate /imdb-citextn<1K0 likes2.4k downloads4y agoHugging Face04uclanlp /OpenVLHarness-Evaluation-Datasets OpenVLHarness evaluation datasets Processed evaluation splits used by OpenVLHarness (project page). Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.imagevisual-question-answering10K<n<100K1 likes1.6k downloads1d agoHugging Face05HAERAE-HUB /K2-EvalResearch Paper coming soon! K2EvalK^{2} EvalK2Eval K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion. Benchmark Overview The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.textn<1K8 likes1.5k downloads2y agoHugging Face06jang1563 /narrow-model-safety-eval Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.tabularothern<1K1 likes980 downloads13d agoHugging Face07CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes660 downloads5mo agoHugging Face08nvidia /PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval VoMP: Predicting Volumetric Mechanical Properties Dataset Description: The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties. We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions. This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.text1K<n<10K3 likes586 downloads8mo agoHugging Face09TechWolf /JobBERT-evaluation-dataset JobBERT evaluation dataset 💾 This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation. The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board. Access the JobBERT paper here: https://arxiv.org/abs/2109.09605 BibTeX Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.texttext-classification10K<n<100K0 likes560 downloads1y agoHugging Face10KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes465 downloads5mo agoHugging Face11prometheus-eval /shoulders-of-giants Shoulders of Giants Can AI agents build on a scientist's work and write a follow-up paper? Each of the 15 tasks gives an agent a published (or recently submitted) scientific paper and a follow-up research direction proposed by the paper's own author. The agent has to carry out the research and write the follow-up paper. Code, rubrics and graders: GitHub. Leaderboard and every judge verdict: website. Subsets Subset Content benchmark/ The 15 tasks:… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/shoulders-of-giants.documentn<1K0 likes446 downloads3d agoHugging Face12RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes420 downloads4y agoHugging Face13allenai /tulu-3-harmbench-evalThis data comes from the HarmBench benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one. textn<1K3 likes412 downloads1y agoHugging Face14ckadirt /mev1_evalstextn<1K0 likes411 downloads3y agoHugging Face15Cross-Mergeability /beetle-merge-eval Beetle merged models — benchmark evaluation against their parents Minimal-pair benchmark accuracy for the Beetle merged models published in the Mergeability org, scored against their own parent models and, where one exists, the jointly-trained ceiling on the same harness. The existing sweep datasets (Mergeability/merge-sweep-results, Mergeability-2/mergeability-results) record merge quality in nats (NLL, delta_floor, rel_damage, barrier, geometry). They contain no downstream… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/beetle-merge-eval.tabular1K<n<10K0 likes390 downloads2mo agoHugging Face16allganize /RAG-Evaluation-Dataset-KO Allganize RAG Leaderboard Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다. RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다. RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다. Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.textn<1K123 likes378 downloads2y agoHugging Face17autoiac-project /iac-eval IaC-Eval dataset (v1.1) IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities. This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now). | Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper | 2. Usage instructions Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.texttext-generationn<1K7 likes365 downloads2y agoHugging Face18FerrariKazu /rhan-eval-sweeptabularn<1K0 likes343 downloads18d agoHugging Face19CharlieLLL /SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923 Fixed Solo350 u355 + MiniMax-M2.7: orchestration cost study Best observed cost tradeoff: compact coordinator decisions plus soft review at the existing hard limit (at most 12 worker turns). M2.7 metered token cost falls 59.3%, while mean solved tasks decrease from 90.00 to 87.67/150. Accuracy equivalence was not established. This closed study contains 6 designs and 16 complete independent runs on the same 150 tasks (2400 scored task/run pairs), each with an independent audit.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923.tabular1K<n<10K0 likes321 downloads17d agoHugging Face20allganize /RAG-Evaluation-Dataset-JA Allganize RAG Leaderboard とは Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.textn<1K42 likes285 downloads2y agoHugging Face21UW-Madison-Lee-Lab /MMLU-Pro-CoT-Eval Dataset Details Modality: Text Format: CSV Size: 100K - 1M rows Total Rows: 248,836 License: MIT Libraries Supported: datasets, pandas, croissant Structure Each row in the dataset includes: question: The query posed in the dataset. answer: The correct response. category: The domain of the question (e.g., math, science). src: The source of the question. id: A unique identifier for each entry. chain_of_thoughts: Step-by-step reasoning steps leading to the answer.… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Eval.text100K<n<1M0 likes281 downloads2y agoHugging Face22AvoCahDoe /llava-15-rlmpq-vlm-eval-results RL-MPQ VLM Evaluation Artifacts Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation. Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results Collections (by base VLM) RL-MPQ VLM — LLaVA-1.5-13B — HF collection RL-MPQ VLM — LLaVA-1.5-7B — HF collection RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection RL-MPQ VLM — Qwen2-VL-7B — HF collection Model repos RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.imagevisual-question-answeringn<1K0 likes274 downloads4mo agoHugging Face23egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face24agentic-moral-alignment /matrix-game-evaltabular10K<n<100K0 likes225 downloads5mo agoHugging Face25RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes217 downloads4y agoHugging Face26felixleungsc /paperswithcode-data-evaluation-tables Process data from paperswithcode See https://huggingface.co/datasets/pwc-archive/files/tree/main. Download and unzip evaluation tables: curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz" gunzip jul-28-evaluation-tables.json.gz Install jq. See https://jqlang.org/. If on Debian/Ubuntu, install with sudo apt-get install jq. Example jq to extract: jq -r ' def process(parent): .task as $current_task | (if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.text100K<n<1M1 likes212 downloads1y agoHugging Face27OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes206 downloads3y agoHugging Face28smksaha /apt-eval 🚨 APT-Eval Dataset 🚨 Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing 📝 Paper, 🖥️ Github, 🎥 Recording This repository contains the official dataset of the ACL 2025 paper 'Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing' APT-Eval is the first and largest dataset to evaluate the AI-text detectors behavior for AI-polished texts. It contains almost 15K text samples, polished by 5 different LLMs, for 6 different domains, with 2 major… See the full description on the dataset page: https://huggingface.co/datasets/smksaha/apt-eval.tabulartext-classification10K<n<100K2 likes199 downloads11mo agoHugging Face29schema-eval /compliance-sycophancy-cot Compliance-Sycophancy CoT Analysis When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating. Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.tabulartext-classificationn<1K0 likes177 downloads1mo agoHugging Face30allenai /olmo-eval-strongrejectThis data comes from the StrongREJECT benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one. Permitted Use The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer This… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-strongreject.text10K<n<100K1 likes176 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.