Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes82k downloads1y agoHugging Face02nz00shuuuu /ltaf-haystack-fixedtabular10K<n<100K0 likes391 downloads6mo agoHugging Face03ameyhengle /Multilingual-Needle-in-a-Haystack Multilingual Needle in a Haystack (MLNeedle) The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.text10K<n<100K3 likes361 downloads1y agoHugging Face04nz00shuuuu /capture24-ts-haystack-cottext10K<n<100K2 likes284 downloads7mo agoHugging Face05nz00shuuuu /uk-dale-haystack UK-DALE-Haystack A controlled additive-needle benchmark for long-context time-series language models built on top of UK-DALE (Kelly & Knottenbelt, 2015), the canonical UK domestic appliance-level + whole-house power demand dataset. Each sample is a 6-second-sampled mains active-power trace with one or more real per-appliance bouts inserted at known locations. A QA prompt asks the model to detect, count, localize, order, or reason about those bouts across five context lengths from 15… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/uk-dale-haystack.tabularquestion-answering10K<n<100K0 likes234 downloads5mo agoHugging Face06tsunghanwu /visual_haystacks Visual Haystacks Dataset Card Dataset details Dataset type: Visual Haystacks (VHs) is a benchmark dataset specifically designed to evaluate the Large Multimodal Model's (LMM's) capability to handle long-context visual information. It can also be viewed as the first vision-centric Needle-In-A-Haystack (NIAH) benchmark dataset. Please also download COCO-2017's training set validation set. Data Preparation and Benchmarking Download the VQA questions:huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/tsunghanwu/visual_haystacks.text10K<n<100K10 likes182 downloads2y agoHugging Face07nz00shuuuu /capture24-ts-haystack-fixed-needle Capture24 TS-Haystack — Fixed Needle Length Long-context retrieval / reasoning benchmark over Capture24 wrist-worn accelerometer recordings, used in Recursive Agents are Effective Time Series Reasoners (ARTS-RLM). This repository supersedes nz00shuuuu/capture24-ts-haystack-cot for the paper's main capture24 experiments. Differences: Fixed (absolute-ms) needle length of 3–10 s across every context length instead of needles that scale with context. With a 7200 s haystack the needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.textquestion-answering10K<n<100K0 likes172 downloads5mo agoHugging Face08nz00shuuuu /sleep_psg_ts_haystacktabular10K<n<100K0 likes156 downloads6mo agoHugging Face09nz00shuuuu /ltaf-haystacktabular10K<n<100K0 likes155 downloads6mo agoHugging Face10Salesforce /summary-of-a-haystack Dataset Card for SummHay This repository contains the data for the experiments in the SummHay paper. Accessing the Data We publicly release the 10 Haystacks (5 in conversational domain, 5 in the news domain). Each example follows the below format: { "topic_id": "ObjectId()", "topic": "", "topic_metadata": {"participants": []}, // can be domain specific "subtopics": [ { "subtopic_id": "ObjectId()", "subtopic_name": ""… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/summary-of-a-haystack.textsummarizationn<1K5 likes117 downloads2y agoHugging Face11brandenchan /haystack-pipelines-no-paramstext1K<n<10K0 likes84 downloads4y agoHugging Face12wenbopan /anti-haystack Dataset Card for "anti-haystack" This dataset contains samples that resemble the "Needle in a haystack" pressure testing. It can be helpful if you want to make your LLM better at finding/locating short facts from long documents. Data Structure Each sample has the following fields: document: A long and noisy reference document which can be a story, code, book, or manual in both English and Chinese (10%). question: A question generated with GPT-4. The answer can always be… See the full description on the dataset page: https://huggingface.co/datasets/wenbopan/anti-haystack.texttext-generation1K<n<10K6 likes83 downloads3y agoHugging Face13jdwiberg /HaystackID_MQP_2025-2026document1K<n<10K0 likes83 downloads11mo agoHugging Face14HaystackBot /medrag-pubmed-chunk-with-embeddingstext10K<n<100K0 likes75 downloads1y agoHugging Face15nz00shuuuu /urbansound-haystack Urban-Sound-Haystack A long-context urban-audio QA benchmark across 10 task types and 4 context lengths (100 s, 15 min, 30 min, 1 h). Each soundscape is synthesised by Scaper from UrbanSound8K foreground events over TUT acoustic-scene backgrounds, sampled at 16 kHz mono PCM_32. Two of the ten tasks (anomaly_detection, anomaly_localization) draw from a parallel pool where every soundscape contains exactly one out-of-vocabulary event from ESC-50 (glass_breaking or crying_baby). This… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/urbansound-haystack.audioaudio-classification10K<n<100K0 likes48 downloads5mo agoHugging Face16vblagoje /haystack-pipelinestext1K<n<10K1 likes43 downloads4y agoHugging Face17RockMan256 /needle-in-a-haystack-ru-48k needle-in-a-haystack-ru-48k Русскоязычный датасет для обучения tool-calling в Home Assistant. Полный вариант (〜48k примеров) в формате "needle in a haystack" — в каждой строке огромный список тулов, модель должна найти нужный. Структура (JSONL, одна запись на строку) { "query": "опусти жалюзи на кухне", "tools": "[{\"name\":\"HassTurnOn\",...}, ...]", "answers": "[{\"name\":\"HassTurnOff\",\"arguments\":{\"name\":\"blinds.kitchen\"}}]" } query — фраза… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/needle-in-a-haystack-ru-48k.text10K<n<100K0 likes37 downloads3mo agoHugging Face18LegionIntel /needle-in-a-haystack-biographies-v2tabular1K<n<10K0 likes28 downloads2y agoHugging Face19tsunghanwu /visual_haystacks_v0 Visual Haystacks Dataset Card Dataset details Dataset type: Visual Haystacks (VHs) is a benchmark dataset specifically designed to evaluate the Large Multimodal Model's (LMM's) capability to handle long-context visual information. It can also be viewed as the first visual-centric Needle-In-A-Haystack (NIAH) benchmark dataset. Please also download COCO-2017's training set validation set. Data Preparation and Benchmarking Download the VQA questions:huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/tsunghanwu/visual_haystacks_v0.text10K<n<100K0 likes23 downloads2y agoHugging Face20RockMan256 /needle-in-a-haystack-lfm-48k needle-in-a-haystack-lfm-48k Датасет для обучения tool-calling в Home Assistant в формате LFM (ready-to-train для LFM2.5 / LFM-семейства). Содержит параллельные русские и английские выборки в одном репозитории. Файлы needle_ru_48k.jsonl — русская выборка needle_en_48k.jsonl — английская выборка Структура (JSONL, одна запись на строку) { "query": "опусти жалюзи на кухне", "tools": "[{\"name\":\"HassTurnOn\",...}, ...]", "answers":… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/needle-in-a-haystack-lfm-48k.text10K<n<100K0 likes22 downloads3mo agoHugging Face21murathankurfali /c2_haystacktext1K<n<10K0 likes18 downloads1y agoHugging Face22Ashima /qwen3_0.6b-task738_augmented_needle_in_a_haystack_Mar16-1507_blendedtabularn<1K0 likes18 downloads7mo agoHugging Face23murathankurfali /c1_haystacktext1K<n<10K0 likes17 downloads1y agoHugging Face24vblagoje /haystack-pipelines-v2text1K<n<10K0 likes16 downloads4y agoHugging Face25Ashima /qwen3_0.6b_needle_in_a_haystack_Mar17-1948_blendedtabularn<1K0 likes15 downloads7mo agoHugging Face26LegionIntel /needle-in-a-haystack-biographies-v0tabular10K<n<100K0 likes14 downloads2y agoHugging Face27vblagoje /haystack-pipelines-v3text1K<n<10K1 likes13 downloads4y agoHugging Face28Ashima /qwen3_0.6b-rlvr_Feb24-1812_datamix_top10_augmented_needle_in_a_haystack_Feb25-0119textn<1K0 likes13 downloads8mo agoHugging Face29Ashima /qwen3_0.6b_needle_in_a_haystack_Mar17-1948tabularn<1K0 likes10 downloads7mo agoHugging Face30arjun3103 /haystack_context_relevance_evaltextn<1K1 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.