datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adaption-financial-qa-with-tables
Financial Filing QA with Reasoning (Augmented)
Numerical questions over company-filing tables and text, answered with a worked reasoning path.
Rows
11,976
Domain
finance
Format
data.parquet, one row per example
Licence
other
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
prompt
The prompt (user turn) as uploaded.
completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-financial-qa-with-tables.table-vqa
Dataset description
The table-vqa Dataset integrates images of tables from the dataset AFTdb (Arxiv Figure Table Database) curated by cmarkea.
This dataset consists of pairs of table images and corresponding LaTeX source code, with each image linked to an average of ten questions and answers. Half of the Q&A pairs are in English and the other half in French. These questions and answers were generated using Gemini 1.5 Pro and Claude 3.5 sonnet, making the dataset well-suited for… See the full description on the dataset page: https://huggingface.co/datasets/cmarkea/table-vqa.TableLLM-SFT
TableLLM-SFT
| Paper | Model | Github | Homepage |
TableLLM-SFT is a training set containing a number of splits on different benchmarks. This training set is used to fine-tuning TableLLM-8b, which are based on Llama3.1-8b-instruct.
needle-in-a-table-pro
NIAT-Pro: Needle-In-A-Table-Pro
NIAT-Pro is a benchmark for evaluating how well large language models understand and reason over large tables under controlled variations of tabular format, table size, and information position.
It extends the original Needle-In-A-Table setting from simple cell lookup to one-hop, two-hop, and four-hop tasks, and studies performance across 11 table representations:
CSV
TSV
PSV
JSON
XML
YAML
Markdown
HTML
LaTeX
SQL
Free-form text
NIAT-Pro is designed… See the full description on the dataset page: https://huggingface.co/datasets/NIAT-Pro/needle-in-a-table-pro.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle
field
value
experiment
One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection).
date_generated
2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.adaption-financial-qa-with-tables-aug
Financial Filing QA with Reasoning (Augmented)
Numerical questions over company-filing tables and text, answered with a worked reasoning path.
Rows
11,976
Domain
finance
Format
data.parquet, one row per example
Licence
other
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-financial-qa-with-tables-aug.2026-08-04-table2-instruction-tuning-9284-filtered-8192
Table 2 instruction-tuning mixture — spec-filtered, 8192-safe (9,284 examples)
The paper's Table 2 instruction-tuning mixture, spec-filtered, with the single row that
cannot fit an 8,192-token window removed. No difficult-advice data — this is the
general instruction-tuning half on its own.
field
value
experiment
Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment and trimmed to fit max_seq_len 8192… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-9284-filtered-8192.md-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.2026-08-04-table2-instruction-tuning-mixture-spec-filtered
Table 2 instruction-tuning mixture, spec-filtered
A reproduction of the paper's Table 2 instruction-tuning mixture at its exact per-source
sample counts, plus an LLM spec-alignment filter and the per-sample judge verdicts, so
the filter can be re-cut at any threshold without paying to re-judge.
field
value
experiment
Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment
date_generated
2026-08-04
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-mixture-spec-filtered.table-sft-eval-predictions
💾 Raw Predictions for "What Really Matters for Table LLMs?"
This dataset contains the raw model outputs from the experiments in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.2026-08-04-qwen36-table2-80-synthdoc-self-reflect-20-sft-bundle
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% training bundle
field
value
experiment
One-epoch Qwen3.6-27B assistant-only LoRA SFT, mixing filtered Table-2 instruction data with first-person SynthDoc self-reflection at 80/20 by loss-bearing tokens.
date_generated
2026-08-04
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md; both upstream corpora connect to this target.
source_repo
Matthew-Bozoukov/teaching_claude_why_replication… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-table2-80-synthdoc-self-reflect-20-sft-bundle.2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture
Table-2 (9,000) + synthdoc difficult-advice (1,000, trait-balanced), all rows <= 8,000 tokens
10,000-example SFT mixture for Qwen3.6-27B. Train on mixture_think.jsonl — every
assistant turn carries a think block, which the trainer's preserve-thinking gate requires.
field
value
experiment
90/10-by-examples SFT mixture: 9,000 spec-filtered Table-2 instruction rows + 1,000 difficult-advice documents drawn evenly across all 9 constitution traits
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture.multimodal_rag_complex_table_extractor_teaser
🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.2026-08-06-table2-9284-synthdoc-716-train
Training bundle — 2026-08-06-table2-9284-synthdoc-716-train
code.tar.gz (trainer, src/, configs/) plus mixture_think.jsonl
(10,000 rows). The pod untars it, copies the jsonl to data/, and runs
configs/train/2026-08-25_lora_qwen36_table2_9284_synthdoc_716.yaml.
field
value
experiment
Table 2 (9,284 spec-filtered) + synthdoc difficult-advice (716, evenly across 9 traits)
date_generated
2026-08-06
constitution
claude_distilled_12_principles_mid — 9 principles; used… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-table2-9284-synthdoc-716-train.2026-08-04-sft-mixture-table2-8000-plus-synthdoc-2203
SFT mixture — Table 2 instruction-tuning (8,000, filtered) + difficult advice (2,203)
10,203 examples. Combines a spec-filtered reproduction of the paper's
Table 2 instruction-tuning mixture with the full difficult-advice corpus generated against a
9-principle distilled constitution.
field
value
experiment
SFT mixture pairing general instruction-tuning data with constitution-aligned difficult-advice data, for the Teaching Claude Why replication
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-sft-mixture-table2-8000-plus-synthdoc-2203.Table-Instructs
📚 Table-Instructs
Bundled instruction-tuning corpora used to train the table LLMs in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
This dataset re-packages the four training corpora used in the paper as a single HF dataset so… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/Table-Instructs.finance-legal-mrc_merged-table
데이터셋 설명
shchoice/finance-legal-mrc 데이터 중 병합된 테이블만 추출한 뒤 이미지와 함께 저장한 데이터입니다.
2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.2026-08-08-table2-9000-synthdoc-1000-trait-balanced-train-mixture
Table-2 (9,000) + synthdoc difficult-advice (1,000, trait-balanced)
10,000-example SFT mixture for Qwen3.6-27B. Train on mixture_think.jsonl — every
assistant turn carries a think block, which the trainer's preserve-thinking gate requires.
field
value
experiment
90/10-by-examples SFT mixture: 9,000 spec-filtered Table-2 instruction rows + 1,000 difficult-advice documents drawn evenly across all 9 constitution traits
date_generated
2026-08-08
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-08-table2-9000-synthdoc-1000-trait-balanced-train-mixture.TableQA-Instruct-v1
TableQA-Instruct-v1
TableQA-Instruct-v1 is a synthetic table question-answering dataset designed for supervised fine-tuning of language models on structured data understanding. It contains table-based QA examples across domains such as environment, health, sports, technology, science, education, finance, business, entertainment, and geography. The dataset helps train models to read tables, answer factual questions, compare values, identify minimum/maximum entries, and perform… See the full description on the dataset page: https://huggingface.co/datasets/kd13/TableQA-Instruct-v1.table-r1-zero-verl
Table-R1-Zero (VERL Format)
This dataset contains 69,265 table reasoning problems from the Table-R1-Zero-Dataset, converted to VERL (Volcano Engine Reinforcement Learning) format for reinforcement learning training workflows.
Source: Table-R1/Table-R1-Zero-Dataset
License: Apache 2.0
Note: System prompts have been removed from all examples for better compatibility with other VERL datasets. The dataset now contains only user messages with table reasoning problems. Ground truth… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/table-r1-zero-verl.relationalrag-tables-artifact-data
RelationalRAG-Tables Artifact Data
This public dataset repository mirrors the lightweight data payloads from the
RelationalRAG-Tables reproducibility artifact:
fixtures/: smoke and synthetic fixtures used by the CPU-only artifact path.
artifacts/: stored JSON outputs and manifest.json hashes for reported
experimental numbers.
The repository does not redistribute heavyweight upstream benchmark corpora
such as BIRD or HybridQA, nor does it redistribute pretrained model weights.… See the full description on the dataset page: https://huggingface.co/datasets/lexuanbach/relationalrag-tables-artifact-data.central-kurdish-correction-table
Central Kurdish Orthographic Correction Table
This repository contains a correction table designed to standardize Central Kurdish (Sorani) script for speech and language technology applications.
Orthographic variation is a major challenge in Kurdish NLP. Different spellings and writing conventions for the same words can introduce inconsistencies in ASR training, evaluation, and downstream NLP systems.
This table provides mappings from non-standard or inconsistent forms to their… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-correction-table.dst-table-prompts-bt
DST Table Prompts
A Danish instruction-tuning dataset of (prompt, article) pairs derived from
Danmarks Statistik (Statistics Denmark) publications.
Each example pairs a natural-language user request — embedding the actual
markdown table — with the real statistician-written article as the target
response. The prompts are LLM-generated and vary in style, tone, and table
placement; the table data and article text come directly from the source
dataset.
Dataset description… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/dst-table-prompts-bt.FinTagging1000_table
FinTagging Table Context Extraction Split
For HTML table inputs, the target is a JSON list of numeric entity/datatype pairs enriched with deterministic row and column context. Missing row or column context is represented as null.
The dataset is derived from FinTagging_800_200_HF and preserves the original
train/test assignment by source_sample_idx and context_id. The XBRL concept
tag is intentionally omitted from the target.
Splits
Split
Samples
Output… See the full description on the dataset page: https://huggingface.co/datasets/lm2445/FinTagging1000_table.gene-r1-go-sft-table2-reconstructed
Gene-R1 GO SFT Table 2 Reconstructed Splits
Small workshop dataset used for tokenizer-transfer experiments with ncbi/Gene-R1-1B.
Rows are reconstructed from released Gene Ontology benchmark materials into Gene-R1-style prompt/completion text for tokenizer transfer experiments:
train.jsonl: 2400 rows, 800 BP + 800 MF + 800 CC
validation.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
test.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
Each row contains metadata plus prompt, completion… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/gene-r1-go-sft-table2-reconstructed.clinical_table_clear_safety_v0.1Clinical Table Clear Safety
PurposeDecide when a clinician must clear irrelevant clutter before acting.
You receive:
table_clutterirrelevant or biasing context
live_evidencecurrent clinical signals
proposed_action
You output one JSON object:
table_clear_requiredyes or no
clear_stepsone sentence describing what to ignore or reset
correct_actionone sentence describing what to do next
Scoring
table_clear_accuracy
clear_steps_similarity
correct_action_similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_table_clear_safety_v0.1.custom_tablesThis table has a similar format as the synthetic_text_to_sql (names the sql_context, sql_prompt, and sql columns).
