datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WikiTableQuestionsTabMWPReasoning-Table
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
The Reasoning-Table dataset is a high-quality, reasoning dataset designed for table reasoning tasks.
📁 Directory Structure
This repository is organized by task. Each subfolder contains task-specific reasoning data, including raw and filtered versions. Here is an overview:
├── fetaqa/
├── feverous/
├── finqa/
├── gsm8k/
├── hitab/
├── hybridqa/
├── multihierttt/
├── ottqa/
├── tabfact/
├── tatqa/
├──… See the full description on the dataset page: https://huggingface.co/datasets/TableQAKit/Reasoning-Table.WikiTableQuestionsSelectionThis dataset is a curated subset of the original WikiTableQuestions (Pasupat and Liang, 2015). To address inconsistencies and inaccuracies found in the source material—such as ambiguous queries and incorrect ground-truth labels—this version consists of 100 hand-selected examples. Each entry has been verified to ensure high data quality and factual alignment, making it an ideal benchmark for precise table-based QA evaluation.
WTQTable-GPT
Table-GPT: Table-tuned GPT for Diverse Table Tasks
This repository contains training and test datasets for the SIGMOD'24 paper Table-GPT: Table-tuned GPT for Diverse Table Tasks. The source code for data generation and task evaluation are available here: https://github.com/microsoft/Table-GPT, which can be used to generate more training data for table-related tasks.
Task Descriptions
We collect (or synthesize) 18 diverse table-related tasks, which are summarized in… See the full description on the dataset page: https://huggingface.co/datasets/LipengCS/Table-GPT.TAT-DQASQAhitabparsebench-table-track
ParseBench Table Track plus Financial Split
This dataset is a mirror of the table dimension of llamaindex/ParseBench, packaged together with a curated Financial Split that we built for evaluating OCR systems on insurance and financial filings.
It contains,
All 503 PDFs of the ParseBench table track.
table.jsonl, the original ground truth (one HTML table per page, plus easy or hard difficulty tag).
financial_split/, our 151 page financial slice plus the 117 dropped non financial… See the full description on the dataset page: https://huggingface.co/datasets/roma2025/parsebench-table-track.TableInstruct
Citation
@misc{wu2024tablebenchcomprehensivecomplexbenchmark,
title={TableBench: A Comprehensive and Complex Benchmark for Table Question Answering},
author={Xianjie Wu and Jian Yang and Linzheng Chai and Ge Zhang and Jiaheng Liu and Xinrun Du and Di Liang and Daixin Shu and Xianfu Cheng and Tianzhen Sun and Guanglin Niu and Tongliang Li and Zhoujun Li},
year={2024},
eprint={2408.09174},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/TableInstruct.TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.adaption-financial-table-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_table_qa
This dataset contains question-answer pairs derived from corporate financial statements, including balance sheets and income statements for various companies. The prompts require extracting specific line items, calculating year-over-year percentage changes, or converting scaled values (millions/thousands) to absolute dollar amounts. Each completion provides the… See the full description on the dataset page: https://huggingface.co/datasets/RaffiArdhi/adaption-financial-table-qa.multihiertt-tablevqamd-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.table-sft-eval-predictions
💾 Raw Predictions for "What Really Matters for Table LLMs?"
This dataset contains the raw model outputs from the experiments in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.fullstackarena-table-study-v1
FullStackArena three-site table study (private working dataset)
Protocol fullstackarena-three-site-table-study-v1 version 2.37 (status qualified). Main table: 4 models × 3 cells × 5 sites = 60 arms, each the whole signed task set of its site: arenagram (132 tasks), rideapp (120 tasks), auction (100 tasks), cross_website (24 tasks), cross_website_deadline (24 tasks); 4,800 scored attempts, max_steps=30, 10 workers on 10 isolated lanes per batch.
Cells:
table1_no_time_capped: no… See the full description on the dataset page: https://huggingface.co/datasets/GroupieSteven/fullstackarena-table-study-v1.2026-08-26-table2-9284-low-stakes-716-train
Low-stakes difficult advice, mixed for training (10,000)
The swap-in twin of LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train. Same 9,284
benign rows, same 716 scenarios, same renderer, same seed — the difficult-advice half
replaced by its low-stakes rewrite. Train this against that control and the only thing
that differs is the magnitude of what the 716 scenarios put at risk.
field
value
experiment
Low-stakes arm of difficult advice: does lowering the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-table2-9284-low-stakes-716-train.FreeformTableQAhistorical-table-reading-order
Reading order in historical documents
A small, reproducible study of one question: on a historical page with more than
one column, the hard part is usually not finding the text lines, but deciding
the order to read them in, and that order is what a text recogniser is
ultimately given.
Live results (every page browsable, side by side): https://valhtrdata01.z1.web.core.windows.net/
Code and data: https://github.com/AbhiPandit1/historical-table-reading-order
What is… See the full description on the dataset page: https://huggingface.co/datasets/abhishekjha1008/historical-table-reading-order.2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.2026-08-04-table2-synthdoc-h200x4-train
Training bundle: Table2 (8,000) + synthdoc (2,202), 4xH200 DDP
code.tar.gz (trainer, src/, configs/) plus mixture_think.jsonl (10,202 rows).
Non-reasoning rows carry the empty think marker as CONTEXT; mask_empty_think: true
excludes those tokens from the loss. synthdoc rows keep real reasoning traces and are
supervised. Config: configs/train/2026-08-25_lora_qwen36_table2_synthdoc.yaml.
2026-08-27-table2-9284-good-ai-fiction-716-train
Table2 9,284 + Good AI Fiction 716 — SFT training mixture
field
value
experiment
The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.2026-08-04-table2-only-9284-h200x4-train
Training bundle: Table 2 only, 9,284 examples (no difficult-advice)
code.tar.gz plus mixture_think.jsonl. Every row carries the empty think marker as
context; mask_empty_think: true keeps those tokens out of the loss.
Config: configs/train/2026-08-25_lora_qwen36_table2_only.yaml.
2026-08-25-table2-9284-difficult-advice-verbose-716-train
Verbose-CoT arm: the published difficult-advice mixture with the 716 difficult-advice reasoning traces expanded ~3x in length, same ideas, to isolate deliberation length from content.
field
value
experiment
Verbose-CoT arm: the published difficult-advice mixture with the 716 difficult-advice reasoning traces expanded ~3x in length, same ideas, to isolate deliberation length from content.
date_generated
20260825
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-716-train.FreeformTableQASelectionThis dataset represents a curated subset of the FreeformTableQA collection, refined to provide a more rigorous benchmark for table-based reasoning. While the original dataset covers a broad range of "free-form" queries—which often include complex, semi-structured, or non-grid layouts—it also contains instances of noise and misaligned labels. To ensure higher evaluation accuracy, this version features 100 manually verified examples where the natural language queries and tabular evidence have… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/FreeformTableQASelection.TableTex
TableTex
Dataset Summary
TableTex is a dataset for table image-to-LaTeX generation, comprising
19,000 renderable LaTeX table annotations extracted from scientific articles
on arXiv. Each annotation corresponds to a table image that can be generated
locally using the provided generate_images.py script. The rendered images
are not included in the repository, which keeps the distributed dataset
compact while providing a reproducible image-generation pipeline.
To… See the full description on the dataset page: https://huggingface.co/datasets/yunfanyang1/TableTex.xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002-details.2026-09-01-table2-9284-difficult-advice-rewritten-702-train-mixture
Table2-9284 + REWRITTEN difficult-advice-702 (thinking SFT mixture)
Rewritten variant of the principle-scoped difficult-advice-702 mixture: the 9,284 standard SFT rows are
byte-identical; the 702 difficult-advice rows have their REASONING and ANSWER rewritten.
field
value
experiment
1-epoch LoRA SFT mixture for Qwen3.6-27B: 9,284 Table2 rows + 702 rewritten difficult-advice rows (7.03% synthetic). Rewrite (on top of the earlier ablation): reasoning OPENING -> neutral… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-table2-9284-difficult-advice-rewritten-702-train-mixture.snowfox-financial-table-data
snowfox-financial-table-data
SnowFox — financial-table extraction training (snowfox_financial_table tasks).
Contents
train.jsonl (1680 rows)
validation.jsonl (162 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for SnowFox (Michael Anthony Falabella).
