datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.multi-context-long-answer-datasetMulti-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.multi_news_long_contextmulti_news_long_context_validationmulti_news_long_context_tokenized_n25722_ctx4096multi_news_long_context_trainmulti_step_reasoning_kazakh_context
🇰🇿 Multi-step Reasoning for Kazakh Context
A high-quality dataset designed for complex reasoning, question answering, and text generation tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
10,981
Total Words (approx.)
6,652,450
Avg. Tokens per Sample
605
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/multi_step_reasoning_kazakh_context.multi_news_long_context_tokenized_n486_ctx4096multi_news_long_context_tokenized_n5757_ctx4096EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-3
Code Apply
Processed EditPackFT-Multi Python, Java, Kotlin, and C splits with fuzzy diff generated using heuristics.
Dataset Preparation
Steps to replicate.
For this version --min_lines_between_chunks=3 was used.
Columns
old_contents the old code
new_contents the new code
fuzzy_diff the code segment extracted from diff between old_contents and new_contents
Example
Diff
from kombu import BrokerConnection
from kombu.common import… See the full description on the dataset page: https://huggingface.co/datasets/kseniasych/EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-3.EditPackFT-Multi-apply-fuzzy-diffs-heuristics_context-3
