Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /evaluation-results@misc{muennighoff2022crosslingual, title={Crosslingual Generalization through Multitask Finetuning}, author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel}, year={2022}, eprint={2211.01786}, archivePrefix={arXiv}, primaryClass={cs.CL} }other100M<n<1B10 likes256k downloads3y agoHugging Face02xiachongfeng /GDP-Val-Evaluation-Submission GDPval Submission Dataset This dataset contains model outputs for GDP-Val evaluation. Dataset Structure data/: Contains the main dataset in Parquet format train-00000-of-00001.parquet: Submission data with model outputs deliverable_files/: Contains generated files for tasks that produce file deliverables Organized by task_id dataset_info.json: Metadata about the dataset Columns task_id: Unique identifier for each task sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.textn<1K0 likes13k downloads1y agoHugging Face03dreamdifferent /vam-cross-evaluation-artifacts0 likes8.5k downloads6h agoHugging Face04sciencialab /grobid-evaluation GROBID End-to-End Evaluation Dataset Reference corpora used for GROBID end-to-end benchmarking of scientific-article structuring. Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/ Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/ Official archive (Zenodo): https://zenodo.org/record/7708580 Dataset summary These are the datasets used for GROBID end-to-end benchmarking, covering: metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.document1K<n<10K1 likes8.2k downloads3mo agoHugging Face05CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes7.1k downloads1y agoHugging Face06sandbagging-games /evaluation_logs Evaluation logs from "Auditing Games for Sandbagging" This dataset provides evaluation transcripts produced for the paper "Auditing Games for Sandbagging". Transcripts are provided in Inspect .eval format, see https://github.com/AI-Safety-Institute/sabotage_games for a guide to viewing them. Dataset Details evaluation_transcripts/handover_evals contains the transcripts provided by the red team to the blue team at the beginning of the main round of the game, showing… See the full description on the dataset page: https://huggingface.co/datasets/sandbagging-games/evaluation_logs.3 likes5.1k downloads9mo agoHugging Face07bigcode /evaluation2 likes3.4k downloads3y agoHugging Face08zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face09Voxel51 /Egocentric_10K_Evaluation Dataset Card for Egocentric_10K_Evaluation This is a FiftyOne dataset with 30000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.imageimage-classification10K<n<100K2 likes3k downloads11mo agoHugging Face10mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K39 likes2.9k downloads3y agoHugging Face11Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 409,710,113 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 988,851,570 rows. This dataset is updated monthly, and was last updated on September 27th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular100M<n<1B33 likes2.8k downloads9d agoHugging Face12MERA-evaluation /MERA MERA (Multimodal Evaluation for Russian-language Architectures) Summary MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language. The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the international arena… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.text10K<n<100K11 likes2.7k downloads2y agoHugging Face13CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.6k downloads4mo agoHugging Face14tsinghua-sigs-robot-lab /VeriLoop-Coder-E1-Evaluation-Evidence VeriLoop Coder-E1 Evaluation Evidence This repository contains the public evaluation-evidence packages referenced by the official VeriLoop Coder-E1 benchmark result files. Model repository: tsinghua-sigs-robot-lab/veriloop-coder-e1 Evidence packages Benchmark Evidence directory DeepSWE veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0 SWE-bench Pro veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0 SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.0 likes1.8k downloads2mo agoHugging Face15fantaxy /user-evaluationsdocumentn<1K0 likes1.7k downloads10mo agoHugging Face16mtec-TUB /GPT-4o-evaluation-biases A database to support the evaluation of gender biases in GPT-4o output The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025). Introduction This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.question-answering10K<n<100K0 likes1.7k downloads2y agoHugging Face17ToiTenBao /hallucination-evaluation-results Hallucination evaluation artifacts Canonical COCO and AMBER detection packages are under detection/<dataset>/models/<model>/probes/<probe>/. See detection/index.json for active and pending packages. Each published package has a manifest with exact checkpoint, cache and test-result paths. Seeds are 0, 42 and 1337. TruthPrInt retains its native validation TPR-at-1%-FPR checkpoint; other detection probes use best pooled validation AUROC. Detection thresholds are validation-best F1.… See the full description on the dataset page: https://huggingface.co/datasets/ToiTenBao/hallucination-evaluation-results.image-text-to-text0 likes1.6k downloads11h agoHugging Face18facebook /ShapeR-Evaluation ShapeR Evaluation Dataset We introduce a new dataset of in-the-wild sequences with paired posed multi-view images, SLAM point clouds, and individually complete 3D shape annotations for 178 objects across 7 diverse scenes. In contrast to existing real-world 3D reconstruction datasets which are either captured in controlled setups or have merged object and background geometries or incomplete shapes, this dataset is designed to capture real-world challenges like occlusions, clutter… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ShapeR-Evaluation.image-to-3dn<1K16 likes1.5k downloads9mo agoHugging Face19tsinghua-sigs-robot-lab /VeriLoop-E2-Evaluation-Evidence VeriLoop E2 Evaluation Evidence Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks. This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results. The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.text-generation0 likes1.4k downloads8d agoHugging Face20happynew111 /haotian_data-GPS-lm-evaluation-harness Language Model Evaluation Harness Latest News 📣 [2025/03] Added support for steering HF models! [2025/02] Added SGLang support! [2024/09] We are prototyping allowing users of LM Evaluation Harness to create and evaluate on text+image multimodal input, text output tasks, and have just added the hf-multimodal and vllm-vlm model types and mmmu task as a prototype feature. We welcome users to try out this in-progress feature and stress-test it for themselves, and suggest… See the full description on the dataset page: https://huggingface.co/datasets/happynew111/haotian_data-GPS-lm-evaluation-harness.0 likes1.4k downloads1y agoHugging Face21amazon /music-off-policy-evaluation-benchmark Music Off-Policy Evaluation Dataset Music Off-Policy Evaluation Dataset is a dataset designed for Off-Policy Evaluation (OPE) research. It contains logged interactions from the home page of Amazon Music. Use cases: Benchmarking OPE estimators Evaluating counterfactual ranking policies offline License Music Off-Policy Evaluation Benchmark © 2026 by Amazon is licensed under Creative Commons Attribution-NonCommercial 4.0 International.… See the full description on the dataset page: https://huggingface.co/datasets/amazon/music-off-policy-evaluation-benchmark.1M<n<10M0 likes1.3k downloads3mo agoHugging Face22VLABench /vlm_evaluation_v1.0 Datacard This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses The dataset structure is as follows: vlm_evaluation_v1.0/ ├── CommenSence/ ├── add_condiment_common_sense/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.image1K<n<10K0 likes1.3k downloads2y agoHugging Face23pengyue-polaron /nyush-galaxea-a1-lingbot-va-real-world-evaluations LingBot-VA on Galaxea A1 — Real-World Evaluations Fruit-placement rollouts and open-loop diagnostics of object grounding, layout generalization, and predicted robot motion. Fruit step-1000: lemon-to-plate rollout in the Official layout. Evidence Scale Real closed-loop rollouts 61 archived; 60 scored Matched base-model controls 9 predictions Post-trained diagnostics 48 full-horizon predictions; 1,211 rolling futures Controlled OOD studies 558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.robotics2 likes1.3k downloads29d agoHugging Face24SZLHOLDINGS /szl-frontier-evaluation-receipts FRONTIER · EVALUATION ARCHIVE Inspect the inputs, model responses and scores behind a small recorded evaluation. Published material Integrity evidence Release boundary Evaluation records File hashes; unsigned HOLD; no promotion Open this run's summary · Review the exact source Scores apply to the recorded cases. This archive grants no execution authority. Checkable language-model evaluation records We publish the inputs, measured responses, scoring results… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-frontier-evaluation-receipts.textn<1K0 likes1.3k downloads1d agoHugging Face25ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K18 likes1.3k downloads1y agoHugging Face26ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes1.2k downloads2y agoHugging Face27pwc-archive /evaluation-tables [!CAUTION] This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025. text1K<n<10K0 likes920 downloads1y agoHugging Face28mypersonalsharingspot11 /evaluation_dataset Data manifest and release plan Smoke-Eval/ contains the held-out evaluation sequences copied from the Windows evaluation workspace. All released inference and metric paths resolve to this directory. The source configurations also refer to two additional datasets that are not required for the selected Smoke-Eval inference check. Dataset Role Expected release contents Current status Rice / Ti-Zed-Dji GRT and refinement training/validation; RGB, radar, and depth inputs… See the full description on the dataset page: https://huggingface.co/datasets/mypersonalsharingspot11/evaluation_dataset.0 likes876 downloads1mo agoHugging Face29Gyubeum /AndroidFlux_Evaluation_Output0 likes843 downloads1mo agoHugging Face30nancyH /token_evaluation0 likes787 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.