datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FactBench
FactBench Leaderboard
VERIFY: A Pipeline for Factuality Evaluation
Language models (LMs) are widely used by an increasing number of users, underscoring the challenge of maintaining factual accuracy across a broad range of topics. We present VERIFY (Verification and Evidence Retrieval for Factuality evaluation), a pipeline to evaluate LMs' factual accuracy in real-world user interactions.
Content Categorization
VERIFY considers the verifiability of LM-generated… See the full description on the dataset page: https://huggingface.co/datasets/launch/FactBench.space-launches
Launches and the satellites they carried, 1957 to 2026
This is a copy of the dataset at mcval.org/launches/data, updated there every Monday and here straight after. Cite it by its DOI (see Citing).
Version 2026-10-07. 7,634 launches from 1957-10-04 to 2026-10-07, and 27,969 payloads, every one linked to the launch that carried it.
Made by McVal for Every launch since Sputnik, a 3D globe of the same data. Licensed CC BY 4.0: use it for anything, with credit (see Citing).… See the full description on the dataset page: https://huggingface.co/datasets/mcval/space-launches.CLASH
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
Paper: CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple PerspectivesContact: leeay@umich.edu
Overview
CLASH (Character perspective-based LLM Assessments in Situations with High-stakes) is a benchmark consisting of 345 long-form, human-written dilemmas spanning high-impact domains.
Each dilemma includes a pair of value-related rationales… See the full description on the dataset page: https://huggingface.co/datasets/launch/CLASH.
