datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egocentric_10K_Evaluation
Dataset Card for Egocentric_10K_Evaluation
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.WEIRD
WEIRD
Описание задачи
WEIRD – это расширенная версия подзадачи бинарной классификации оригинального английского бенчмарка WHOOPS!. Датасет оценивает, способна ли мультимодальная модель обнаруживать нарушения здравого смысла в изображениях. Здесь нарушение здравого смысла – это ситуации, противоречащие типичным нормам реальности. Например, пингвины не могут летать, дети не водят автомобили, посетители не накладывают еду официантам, и так далее. В датасете поровну… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/WEIRD.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.evaluation-dataset
DeepSafe Evaluation Dataset
Evaluation set for DeepSafe,
a deepfake detection benchmark.
Tiers
Tier
Samples
Generators
Size
Use
master_eval_small/
198
116
1.7 GB
smoke test, under 2 min
master_eval/
15,454
411
10 GB
the standard benchmark
master_eval_full/
45,954
411
25 GB
complete set
Medium tier composition: 9,954 image, 3,500 audio, 2,000 video.
from huggingface_hub import snapshot_download
snapshot_download("deepsafe/evaluation-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/deepsafe/evaluation-dataset.medical-imaging-model-evaluation-benchmark
医学影像多模态模型评测集(精选示例版)
这是一个面向医学多模态大模型的高质量影像评测集,专门测试模型能否把“看见影像”进一步转化为可解释、可复核、符合临床语境的判断与表达。数据将医学影像与患者描述、病史摘要、检查信息或结构化临床资料配对,覆盖从影像分类、报告生成,到鉴别诊断、治疗方案和胸片质量控制的完整评测链路。
本次公开版本从 2026-07-22 质检通过产物中整理而来,按每个子集最多 50 题进行分层抽样;题量不足 50 的影像质量控制子集完整保留。因此,公开版本包含 5 个任务子集、213 题和 455 个配套影像文件,适合作为医学视觉语言模型的快速对比集、回归测试集和研究教学样例。
数据集亮点
多模态对齐:每条样例同时提供影像和结构化的 question、answer、explanation,支持检查视觉理解、临床语义整合与解释质量。
任务覆盖完整:从“影像是什么”到“如何描述、如何鉴别、如何处置”,并加入真实影像工作流中的胸片质量控制任务。
影像类型丰富:覆盖 X… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-imaging-model-evaluation-benchmark.evaluation
Skill-Aligned Annotation for Text-to-Image Evaluation
Companion dataset for the NeurIPS 2026 paper "Towards Objective Evaluation".
The dataset contains generated images from 7 text-to-image models, evaluated
by 6 human annotators (anonymized) plus an LLM judge across 9 skill-aligned
annotation strategies.
Configs
Config
Rows
Description
images
621
Generated images (621 WebP) with embedded bytes; one row per (prompt_id, generator).
prompts
179
Per-prompt… See the full description on the dataset page: https://huggingface.co/datasets/Skill-Aigned/evaluation.
