Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K17 likes247 downloads3mo agoHugging Face02nbtpj /multi-context-long-answer-datasettext1M<n<10M13 likes214 downloads4y agoHugging Face03TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes183 downloads1y agoHugging Face04kothasuhas /multi_news_long_contexttext10K<n<100K1 likes21 downloads10mo agoHugging Face05kothasuhas /multi_news_long_context_validationtext1K<n<10K0 likes14 downloads10mo agoHugging Face06kothasuhas /multi_news_long_context_tokenized_n25722_ctx409610K<n<100K0 likes13 downloads10mo agoHugging Face07kothasuhas /multi_news_long_context_traintext100K<n<1M0 likes13 downloads10mo agoHugging Face08farabi-lab /multi_step_reasoning_kazakh_contextgated 🇰🇿 Multi-step Reasoning for Kazakh Context A high-quality dataset designed for complex reasoning, question answering, and text generation tasks in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 10,981 Total Words (approx.) 6,652,450 Avg. Tokens per Sample 605 Word Count Distribution (Per Field) The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/multi_step_reasoning_kazakh_context.textquestion-answering10K<n<100K0 likes13 downloads2mo agoHugging Face09kothasuhas /multi_news_long_context_tokenized_n486_ctx4096n<1K0 likes11 downloads10mo agoHugging Face10kothasuhas /multi_news_long_context_tokenized_n5757_ctx40961K<n<10K0 likes9 downloads10mo agoHugging Face11kseniasych /EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-3 Code Apply Processed EditPackFT-Multi Python, Java, Kotlin, and C splits with fuzzy diff generated using heuristics. Dataset Preparation Steps to replicate. For this version --min_lines_between_chunks=3 was used. Columns old_contents the old code new_contents the new code fuzzy_diff the code segment extracted from diff between old_contents and new_contents Example Diff from kombu import BrokerConnection from kombu.common import… See the full description on the dataset page: https://huggingface.co/datasets/kseniasych/EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-3.text10K<n<100K0 likes8 downloads1y agoHugging Face12ksych /EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-30 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.