Team Ai
Datasetpublic

AmazonScience/ListQA

ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models Project page: listqa.github.io · Venue: NeurIPS 2026, Evaluations and Datasets Track ListQA tests whether an LLM can recall many facts from its parameters and arrange them in a list. Most factual QA benchmarks ask for a single answer. ListQA instead has 9,045 human-written, cross-validated questions, and each answer is a list of 3–10 elements. Every question and answer is… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ListQA.

sourceHugging Facecc-by-nc-4.0updated 8d agoView on Hugging Face
0likes151downloads
8 commits on main
92e43ea8d ago

Upload README.md

maykul
ae674b39d ago

Upload README.md

maykul
38c430c9d ago

Add benchmark canary GUID to dataset card

maykul
a3f910c9d ago

Upload README.md

maykul
dd4015c9d ago

Restructure into configs, add dataset card

maykul
d9b902e9d ago

Create README.md

maykul
28e57619d ago

Upload 7 files

maykul
b9c82909d ago

initial commit

maykul