datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.MultilingualMultiModalClassification
Additional Information
To load the dataset,
import datasets
ds = datasets.load_dataset("AmazonScience/MultilingualMultiModalClassification", data_dir="wiki-doc-ar-merged")
print(ds)
DatasetDict({
train: Dataset({
features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'],
num_rows: 8129
})
validation: Dataset({
features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'],
num_rows: 1742
})
test: Dataset({
features:… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/MultilingualMultiModalClassification.massive MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations
for the Natural Language Understanding tasks of intent prediction and slot annotation.
Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing
the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench.migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.SWE-PolyBench_Verified
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is:
Javascript: 100
Typescript: 100
Python: 113
Java: 69
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.bold
Dataset Card for Bias in Open-ended Language Generation Dataset (BOLD)
Dataset Description
Bias in Open-ended Language Generation Dataset (BOLD) is a dataset to evaluate fairness in open-ended language generation in English language. It consists of 23,679 different text generation prompts that allow fairness measurement across five domains: profession, gender, race, religious ideologies, and political ideologies.
Some examples of prompts in BOLD are as follows:
Many… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/bold.SWE-PolyBench_500
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_500.RobustAD
RobustAD Dataset
About the Dataset
RobustAD, specifically designed to evaluate the robustness of anomaly detection models in real-world scenarios. RobustAD features a curated dataset of defect detection images with meticulously controlled distribution shifts across multiple dimensions relevant to practical applications and more closely mirrors real-world deployment scenarios.
RobustAD is designed to cover inspection challenges across multiple industries to ensure the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/RobustAD.mintaka Mintaka is a complex, natural, and multilingual dataset designed for experimenting with end-to-end
question-answering models. Mintaka is composed of 20,000 question-answer pairs collected in English,
annotated with Wikidata entities, and translated into Arabic, French, German, Hindi, Italian,
Japanese, Portuguese, and Spanish for a total of 180,000 samples.
Mintaka includes 8 types of complex questions, including superlative, intersection, and multi-hop questions,
which were naturally elicited from crowd workers.FalseReject
FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models
FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts.
FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.SpIDER-Bench
SpIDER-Bench
Repository dependency graphs for software issue localization — the graph data behind
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
(arXiv:2512.16956).
Each benchmark instance gets one directed multigraph of its repository at the commit the
issue was filed against. Nodes are directories, files, classes and functions carrying
their source; edges are contains / imports / inherits / invokes relations between
them. SpIDER uses these… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SpIDER-Bench.DocTalk
📁 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
➤ 📖 Paper Link
DocTalk is a large-scale, synthetic dialogue corpus created via a three-stage pipeline that converts clusters of related Wikipedia documents into multi-turn, multi-topic information-seeking conversations.
The pipeline comprises:
Document Graph Construction: Sampling up to three related Wikipedia articles per anchor document via a weighted random walk on a directed… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/DocTalk.migration-bench-java-selected
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.asnq
Dataset Card for "asnq"
Dataset Summary
ASNQ is a dataset for answer sentence selection derived from
Google's Natural Questions (NQ) dataset (Kwiatkowski et al. 2019).
Each example contains a question, candidate sentence, label indicating whether or not
the sentence answers the question, and two additional features --
sentence_in_long_answer and short_answer_in_sentence indicating whether ot not the
candidate sentence is contained in the long_answer and if the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/asnq.massive-agentstydi-as2
TyDi-AS2
Dataset Summary
TyDi-AS2 and Xtr-TyDi-AS2 are multilingual Answer Sentence Selection (AS2) datasets comprising 8 diverse languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages.
Both the datasets were created from TyDi-QA, a multilingual question-answering dataset. TyDi-AS2 was created by converting the QA instances in TyDi-QA to AS2 instances (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/tydi-as2.WikiDT
WikiDT: Wikipedia Table Document dataset for table extraction and visual question answering
Dataset Summary
The WikiDT contains multi-level annotations and labels for the question-answering task based on images. Meanwhile, as the questions are answered from some table on the image, and WikiDT provides the table annotation to facilitate the diagnosis of the models and decompose the problem, WikiDT can be also directly used as a
table recognition dataset.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/WikiDT.PersonaLens
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/PersonaLens.migration-bench-java-utg
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-utg.GaRAGe
GaRAGe
A Benchmark with Grounding Annotations for RAG Evaluation
This repository contains the data for the paper (ACL 2025 Findings): GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation.
GaRAGe is a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can identify relevant grounding when generating RAG answers. This benchmark contains 2366 questions of diverse… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/GaRAGe.XRAG
XRAG
1. 📖 Overview
XRAG is a benchmark dataset for evaluating LLMs' generation capabilities in a cross-lingual RAG setting, where questions and retrieved documents are in different languages. It covers two different cross-lingual RAG scenarios:
Cross-lingual RAG with Monolingual Retrieval, where questions are non-English while the retrieved documents are in EnglishCross-lingual RAG with Multilingual Retrieval, where questions are non-English while the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/XRAG.xtr-wiki_qa
Xtr-WikiQA
Dataset Summary
Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages.
This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face).
For translations, we used Amazon Translate.
Languages
Arabic (ar)
Spanish (es)
French (fr)
German (de)
Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.Multi-IaC-Eval
Multi-IaC-Eval
We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation.
Cloudformation: 263
Terraform: 446
CDK (Python): 64
CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.ListQA
ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models
Project page: listqa.github.io · Venue: NeurIPS 2026, Evaluations and Datasets Track
ListQA tests whether an LLM can recall many facts from its parameters and arrange them in a list. Most factual QA benchmarks ask for a single answer. ListQA instead has 9,045 human-written, cross-validated questions, and each answer is a list of 3–10 elements. Every question and answer is… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ListQA.mxevalA collection of execution-based multi-lingual benchmark for code generation.AIDSAFE
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
This dataset contains policy-embedded Chain-of-Thought (CoT) data generated using the AIDSAFE (Agentic Iterative Deliberation for SAFEty Reasoning) framework to improve safety reasoning in Large Language Models (LLMs).
Dataset Overview
Dataset Description
The AIDSAFE Policy-Embedded CoT Dataset is a collection of high-quality, safety-focused Chain-of-Thought (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/AIDSAFE.TISER
TISER
Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models
This repository contains the data for the paper (ACL 2025 Main): Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models.
TISER incorporates a multi-stage inference pipeline that combines explicit reasoning, timeline construction, and iterative self-reflection. The key idea behind our approach is to empower LLMs… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/TISER.TANGO
Dataset Card for TANGO
TANGO (Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation) is a dataset that consists of two sets of prompts to evaluate gender non-affirmative language in open
language generation (OLG).
Intended Use
TANGO is intended to help assess the extent to which models reflect undesirable societal biases relating to the Transgender and Non-Binary (TGNB) community, with the goal of promoting fairness and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/TANGO.ESRRSim
ESRRSim Generated Benchmark
Evaluation benchmark for Emergent Strategic Reasoning Risks (ESRRs) in large language models, generated by the ESRRSim agentic framework. This dataset provides 1,052 evaluation scenarios with paired dual rubrics for assessing both model responses and reasoning traces.
📄 Paper: Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
Dataset Description
⚠️ Disclaimer: All names, organizations, characters, and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ESRRSim.
