Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes86k downloads1y agoHugging Face02AmazonScience /MultilingualMultiModalClassification Additional Information To load the dataset, import datasets ds = datasets.load_dataset("AmazonScience/MultilingualMultiModalClassification", data_dir="wiki-doc-ar-merged") print(ds) DatasetDict({ train: Dataset({ features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'], num_rows: 8129 }) validation: Dataset({ features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'], num_rows: 1742 }) test: Dataset({ features:… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/MultilingualMultiModalClassification.2 likes61k downloads2y agoHugging Face03AmazonScience /massive MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation. Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.text-classification100K<n<1M101 likes16k downloads4y agoHugging Face04AmazonScience /SWE-PolyBench SWE-PolyBench SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is: Javascript: 1017 Typescript: 729 Python: 199 Java: 165 Datasets There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench.tabular1K<n<10K5 likes6.7k downloads1y agoHugging Face05AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes3.9k downloads1y agoHugging Face06AmazonScience /SWE-PolyBench_Verified SWE-PolyBench SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is: Javascript: 100 Typescript: 100 Python: 113 Java: 69 Datasets There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.tabularn<1K5 likes2.6k downloads10mo agoHugging Face07AmazonScience /bold Dataset Card for Bias in Open-ended Language Generation Dataset (BOLD) Dataset Description Bias in Open-ended Language Generation Dataset (BOLD) is a dataset to evaluate fairness in open-ended language generation in English language. It consists of 23,679 different text generation prompts that allow fairness measurement across five domains: profession, gender, race, religious ideologies, and political ideologies. Some examples of prompts in BOLD are as follows: Many… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/bold.texttext-generation1K<n<10K20 likes2k downloads4y agoHugging Face08AmazonScience /SWE-PolyBench_500 SWE-PolyBench SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is: Javascript: 1017 Typescript: 729 Python: 199 Java: 165 Datasets There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_500.tabularn<1K3 likes1.6k downloads1y agoHugging Face09AmazonScience /RobustAD RobustAD Dataset About the Dataset RobustAD, specifically designed to evaluate the robustness of anomaly detection models in real-world scenarios. RobustAD features a curated dataset of defect detection images with meticulously controlled distribution shifts across multiple dimensions relevant to practical applications and more closely mirrors real-world deployment scenarios. RobustAD is designed to cover inspection challenges across multiple industries to ensure the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/RobustAD.image1K<n<10K1 likes1.6k downloads1y agoHugging Face10AmazonScience /mintaka Mintaka is a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models. Mintaka is composed of 20,000 question-answer pairs collected in English, annotated with Wikidata entities, and translated into Arabic, French, German, Hindi, Italian, Japanese, Portuguese, and Spanish for a total of 180,000 samples. Mintaka includes 8 types of complex questions, including superlative, intersection, and multi-hop questions, which were naturally elicited from crowd workers.textquestion-answering100K<n<1M12 likes992 downloads4y agoHugging Face11AmazonScience /FalseReject FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts. FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.texttext-generation10K<n<100K36 likes935 downloads1y agoHugging Face12AmazonScience /SpIDER-Bench SpIDER-Bench Repository dependency graphs for software issue localization — the graph data behind SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization (arXiv:2512.16956). Each benchmark instance gets one directed multigraph of its repository at the commit the issue was filed against. Nodes are directories, files, classes and functions carrying their source; edges are contains / imports / inherits / invokes relations between them. SpIDER uses these… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SpIDER-Bench.tabularfeature-extraction100M<n<1B2 likes722 downloads1mo agoHugging Face13AmazonScience /DocTalk 📁 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities ➤ 📖 Paper Link DocTalk is a large-scale, synthetic dialogue corpus created via a three-stage pipeline that converts clusters of related Wikipedia documents into multi-turn, multi-topic information-seeking conversations. The pipeline comprises: Document Graph Construction: Sampling up to three related Wikipedia articles per anchor document via a weighted random walk on a directed… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/DocTalk.textquestion-answering100K<n<1M2 likes622 downloads1y agoHugging Face14AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes582 downloads1y agoHugging Face15AmazonScience /asnq Dataset Card for "asnq" Dataset Summary ASNQ is a dataset for answer sentence selection derived from Google's Natural Questions (NQ) dataset (Kwiatkowski et al. 2019). Each example contains a question, candidate sentence, label indicating whether or not the sentence answers the question, and two additional features -- sentence_in_long_answer and short_answer_in_sentence indicating whether ot not the candidate sentence is contained in the long_answer and if the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/asnq.textmultiple-choice10M<n<100M2 likes457 downloads3y agoHugging Face16AmazonScience /massive-agents2 likes349 downloads5mo agoHugging Face17AmazonScience /tydi-as2 TyDi-AS2 Dataset Summary TyDi-AS2 and Xtr-TyDi-AS2 are multilingual Answer Sentence Selection (AS2) datasets comprising 8 diverse languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages. Both the datasets were created from TyDi-QA, a multilingual question-answering dataset. TyDi-AS2 was created by converting the QA instances in TyDi-QA to AS2 instances (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/tydi-as2.textquestion-answering10M<n<100M1 likes319 downloads3y agoHugging Face18AmazonScience /WikiDT WikiDT: Wikipedia Table Document dataset for table extraction and visual question answering Dataset Summary The WikiDT contains multi-level annotations and labels for the question-answering task based on images. Meanwhile, as the questions are answered from some table on the image, and WikiDT provides the table annotation to facilitate the diagnosis of the models and decompose the problem, WikiDT can be also directly used as a table recognition dataset. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/WikiDT.table-question-answering100K<n<1M1 likes300 downloads3y agoHugging Face19AmazonScience /PersonaLens PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/PersonaLens.text-generation1K<n<10K2 likes249 downloads1y agoHugging Face20AmazonScience /migration-bench-java-utg MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-utg.texttext-generation1K<n<10K4 likes221 downloads1y agoHugging Face21AmazonScience /GaRAGe GaRAGe A Benchmark with Grounding Annotations for RAG Evaluation This repository contains the data for the paper (ACL 2025 Findings): GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation. GaRAGe is a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can identify relevant grounding when generating RAG answers. This benchmark contains 2366 questions of diverse… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/GaRAGe.text1K<n<10K1 likes221 downloads1y agoHugging Face22AmazonScience /XRAG XRAG 1. 📖 Overview XRAG is a benchmark dataset for evaluating LLMs' generation capabilities in a cross-lingual RAG setting, where questions and retrieved documents are in different languages. It covers two different cross-lingual RAG scenarios: Cross-lingual RAG with Monolingual Retrieval, where questions are non-English while the retrieved documents are in EnglishCross-lingual RAG with Multilingual Retrieval, where questions are non-English while the… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/XRAG.question-answering1K<n<10K4 likes212 downloads1y agoHugging Face23AmazonScience /xtr-wiki_qa Xtr-WikiQA Dataset Summary Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages. This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face). For translations, we used Amazon Translate. Languages Arabic (ar) Spanish (es) French (fr) German (de) Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.textquestion-answering100K<n<1M5 likes203 downloads3y agoHugging Face24AmazonScience /Multi-IaC-Eval Multi-IaC-Eval We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation. Cloudformation: 263 Terraform: 446 CDK (Python): 64 CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.texttext-generationn<1K0 likes116 downloads1y agoHugging Face25AmazonScience /ListQA ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models Project page: listqa.github.io · Venue: NeurIPS 2026, Evaluations and Datasets Track ListQA tests whether an LLM can recall many facts from its parameters and arrange them in a list. Most factual QA benchmarks ask for a single answer. ListQA instead has 9,045 human-written, cross-validated questions, and each answer is a list of 3–10 elements. Every question and answer is… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ListQA.textquestion-answering10K<n<100K0 likes91 downloads3d agoHugging Face26AmazonScience /mxevalA collection of execution-based multi-lingual benchmark for code generation.text-generation0 likes52 downloads2y agoHugging Face27AmazonScience /AIDSAFE Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation This dataset contains policy-embedded Chain-of-Thought (CoT) data generated using the AIDSAFE (Agentic Iterative Deliberation for SAFEty Reasoning) framework to improve safety reasoning in Large Language Models (LLMs). Dataset Overview Dataset Description The AIDSAFE Policy-Embedded CoT Dataset is a collection of high-quality, safety-focused Chain-of-Thought (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/AIDSAFE.texttext-generation10K<n<100K2 likes52 downloads1y agoHugging Face28AmazonScience /TISER TISER Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models This repository contains the data for the paper (ACL 2025 Main): Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models. TISER incorporates a multi-stage inference pipeline that combines explicit reasoning, timeline construction, and iterative self-reflection. The key idea behind our approach is to empower LLMs… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/TISER.textquestion-answering10K<n<100K2 likes46 downloads1y agoHugging Face29AmazonScience /TANGO Dataset Card for TANGO TANGO (Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation) is a dataset that consists of two sets of prompts to evaluate gender non-affirmative language in open language generation (OLG). Intended Use TANGO is intended to help assess the extent to which models reflect undesirable societal biases relating to the Transgender and Non-Binary (TGNB) community, with the goal of promoting fairness and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/TANGO.text-generation1M<n<10M2 likes40 downloads3y agoHugging Face30AmazonScience /ESRRSim ESRRSim Generated Benchmark Evaluation benchmark for Emergent Strategic Reasoning Risks (ESRRs) in large language models, generated by the ESRRSim agentic framework. This dataset provides 1,052 evaluation scenarios with paired dual rubrics for assessing both model responses and reasoning traces. 📄 Paper: Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework Dataset Description ⚠️ Disclaimer: All names, organizations, characters, and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/ESRRSim.texttext-generation1K<n<10K1 likes30 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.