datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.document-haystack-10pages
Dataset Card for document-haystack-10pages
This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/document-haystack-10pages")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.ltaf-haystack-fixedHaystackCraft@article{li2025haystack,
title={Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation},
author={Mufei Li and Dongqi Fu and Limei Wang and Si Zhang and Hanqing Zeng and Kaan Sancak and Ruizhong Qiu and Haoyu Wang and Xiaoxin He and Xavier Bresson and Yinglong Xia and Chonglin Sun and Pan Li},
journal={arXiv preprint arXiv:2510.07414},
year={2025}
}
Multilingual-Needle-in-a-Haystack
Multilingual Needle in a Haystack (MLNeedle)
The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.capture24-ts-haystack-cotDocument_Haystacksuk-dale-haystack
UK-DALE-Haystack
A controlled additive-needle benchmark for long-context time-series language
models built on top of UK-DALE (Kelly & Knottenbelt, 2015), the canonical
UK domestic appliance-level + whole-house power demand dataset.
Each sample is a 6-second-sampled mains active-power trace with one or more
real per-appliance bouts inserted at known locations. A QA prompt asks the
model to detect, count, localize, order, or reason about those bouts across
five context lengths from 15… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/uk-dale-haystack.sleep_psg_ts_haystackcapture24-ts-haystack-fixed-needle
Capture24 TS-Haystack — Fixed Needle Length
Long-context retrieval / reasoning benchmark over Capture24 wrist-worn
accelerometer recordings, used in Recursive Agents are Effective Time Series
Reasoners (ARTS-RLM).
This repository supersedes
nz00shuuuu/capture24-ts-haystack-cot
for the paper's main capture24 experiments. Differences:
Fixed (absolute-ms) needle length of 3–10 s across every context length
instead of needles that scale with context. With a 7200 s haystack the
needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.ltaf-haystackvisual_haystacks
Visual Haystacks Dataset Card
Dataset details
Dataset type: Visual Haystacks (VHs) is a benchmark dataset specifically designed to evaluate the Large Multimodal Model's (LMM's) capability to handle long-context visual information. It can also be viewed as the first vision-centric Needle-In-A-Haystack (NIAH) benchmark dataset. Please also download COCO-2017's training set validation set.
Data Preparation and Benchmarking
Download the VQA questions:huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/tsunghanwu/visual_haystacks.Little_haystack_20251203_20260227HaystackID_MQP_2025-2026summary-of-a-haystack
Dataset Card for SummHay
This repository contains the data for the experiments in the SummHay paper.
Accessing the Data
We publicly release the 10 Haystacks (5 in conversational domain, 5 in the news domain). Each example follows the below format:
{
"topic_id": "ObjectId()",
"topic": "",
"topic_metadata": {"participants": []}, // can be domain specific
"subtopics": [
{
"subtopic_id": "ObjectId()",
"subtopic_name": ""… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/summary-of-a-haystack.haystack-pipelines-no-paramsmedrag-pubmed-chunk-with-embeddingsanti-haystack
Dataset Card for "anti-haystack"
This dataset contains samples that resemble the "Needle in a haystack" pressure testing. It can be helpful if you want to make your LLM better at finding/locating short facts from long documents.
Data Structure
Each sample has the following fields:
document: A long and noisy reference document which can be a story, code, book, or manual in both English and Chinese (10%).
question: A question generated with GPT-4. The answer can always be… See the full description on the dataset page: https://huggingface.co/datasets/wenbopan/anti-haystack.haystack-pipelinesurbansound-haystack
Urban-Sound-Haystack
A long-context urban-audio QA benchmark across 10 task types and
4 context lengths (100 s, 15 min, 30 min, 1 h). Each soundscape is
synthesised by Scaper from
UrbanSound8K foreground events over TUT acoustic-scene backgrounds, sampled
at 16 kHz mono PCM_32. Two of the ten tasks
(anomaly_detection, anomaly_localization) draw from a parallel pool
where every soundscape contains exactly one out-of-vocabulary event from
ESC-50 (glass_breaking or crying_baby).
This… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/urbansound-haystack.needle-in-a-haystack-ru-48k
needle-in-a-haystack-ru-48k
Русскоязычный датасет для обучения tool-calling в Home Assistant. Полный
вариант (〜48k примеров) в формате "needle in a haystack" — в каждой строке
огромный список тулов, модель должна найти нужный.
Структура (JSONL, одна запись на строку)
{
"query": "опусти жалюзи на кухне",
"tools": "[{\"name\":\"HassTurnOn\",...}, ...]",
"answers": "[{\"name\":\"HassTurnOff\",\"arguments\":{\"name\":\"blinds.kitchen\"}}]"
}
query — фраза… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/needle-in-a-haystack-ru-48k.needle-in-a-haystack-biographies-v2visual_haystacks_v0
Visual Haystacks Dataset Card
Dataset details
Dataset type: Visual Haystacks (VHs) is a benchmark dataset specifically designed to evaluate the Large Multimodal Model's (LMM's) capability to handle long-context visual information. It can also be viewed as the first visual-centric Needle-In-A-Haystack (NIAH) benchmark dataset. Please also download COCO-2017's training set validation set.
Data Preparation and Benchmarking
Download the VQA questions:huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/tsunghanwu/visual_haystacks_v0.needle-in-a-haystack-lfm-48k
needle-in-a-haystack-lfm-48k
Датасет для обучения tool-calling в Home Assistant в формате LFM
(ready-to-train для LFM2.5 / LFM-семейства). Содержит параллельные
русские и английские выборки в одном репозитории.
Файлы
needle_ru_48k.jsonl — русская выборка
needle_en_48k.jsonl — английская выборка
Структура (JSONL, одна запись на строку)
{
"query": "опусти жалюзи на кухне",
"tools": "[{\"name\":\"HassTurnOn\",...}, ...]",
"answers":… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/needle-in-a-haystack-lfm-48k.haystack-pipelines-v2c2_haystackv-niah-haystackqwen3_0.6b-task738_augmented_needle_in_a_haystack_Mar16-1507_blendedc1_haystackHaystackCraft
