datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Embodied-Agent-Arena
Embodied Agent Arena — task data
Task data for the Embodied Agent Arena evaluation code.
This task collection contains 1,000 uniquely identified evaluation cases:
W1 370, W2 220, W3 157, W4 183, and W5 70.
The 660 W1/W2/W5 cases use fixed image/video inputs; the 340 W3/W4 cases require
interactive benchmark environments. The task index contains 41 adapter IDs;
several adapters correspond to task types from the same source dataset.
W4 contains thirteen reporting groups over… See the full description on the dataset page: https://huggingface.co/datasets/uuu-Quant/Embodied-Agent-Arena.anonymous_dataset
KnowVis: A Dual-View Benchmark for Diagnosing World-Knowledge Grounding in Text-to-Image Models
📖 Overview
Text-to-image (T2I) models have made substantial progress in visual realism, aesthetic quality, and instruction following. However, real-world prompts often go beyond explicit visual descriptions and require implicit facts, structured knowledge, and domain-specific commonsense. Existing evaluations mainly focus on explicit prompt-to-image semantic alignment or… See the full description on the dataset page: https://huggingface.co/datasets/QuantumWhisper42/anonymous_dataset.
