AmazonScience/ListQA
ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models Project page: listqa.github.io · Venue: NeurIPS 2026, Evaluations and Datasets Track ListQA tests whether an LLM can recall many facts from its parameters and arrange them in a list. Most factual QA benchmarks ask for a single answer. ListQA instead has 9,045 human-written, cross-validated questions, and each answer is a list of 3–10 elements. Every question and answer is… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ListQA.
091
