datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.tabularbenchmachinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.CausalArena
CausalArena public release
This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices.
What is included
scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports.
scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.TALENT
Benchmark Dataset for "A Closer Look at Deep Learning on Tabular Data"
This repository contains supplemental datasets for the paper "A Closer Look at Deep Learning on Tabular Data". The datasets are provided in two zip files: benchmark_dataset.zip and training_dynamic_informations.zip.
1. benchmark_dataset.zip
This zip file contains all the tabular datasets used in the paper. Each dataset is stored in a separate subfolder named after the dataset. Each dataset folder… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/TALENT.p2-etf-tabular-alpha-decay-resultstcga-luad-tabular-open
TCGA-LUAD — Tabular (Open Access)
Open-access TCGA-LUAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:06:47 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-luad-tabular-open.tcga-coad-tabular-open
TCGA-COAD — Tabular (Open Access)
Open-access TCGA-COAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:52:16 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-coad-tabular-open.house-prices-tabular
House Prices - Tabular (with baseline metrics)
A curated, ready-to-use copy of the Kaggle House Prices: Advanced
Regression Techniques training set (1,460 × 81), accompanied by baseline
metrics from an 8-architecture tabular modelling pipeline so newcomers
have a reference point to compare against.
This is the same data as Kaggle's train.csv, sourced from
OpenML id 42165 (canonical mirror).
At a glance
Field
Value
Rows
1,460
Columns
81 (80 features + 1… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/house-prices-tabular.tcga-stad-tabular-open
TCGA-STAD — Tabular (Open Access)
Open-access TCGA-STAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:19:42 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-stad-tabular-open.TabularMath
📊 TabularMath
TabularMath is a tabular mathematical reasoning benchmark introduced in TabularMath: Understanding Math Reasoning over Tables with Large Language Models. It is built via AUTOT2T, a neuro-symbolic pipeline that automatically transforms math word problems into verified tabular reasoning tasks, enabling scalable evaluation without manual table annotation.
TabularMath jointly assesses reasoning accuracy, information retrieval over complex table structures, and… See the full description on the dataset page: https://huggingface.co/datasets/kevin715/TabularMath.demo-tabular-benchmark-containers
📦 Carla HQ — Tabular Benchmark CuratedContainers
Centralized repository of Data Foundry CuratedContainers curated for Carla HQ, TabICLv2, and the next generation of Tabular Foundation Models (TabPFN, EXAONE, Google TabFM).
Each container directory provides:
Columnar Parquet Data (dataset.parquet): Clean, type-normalized, and validated tabular dataset binary.
Standardized Task Molds (task_metadata.predictive-ml-task-mold-v1.json): Problem definitions, target attributes, and… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmark-containers.tcga-lihc-tabular-open
TCGA-LIHC — Tabular (Open Access)
Open-access TCGA-LIHC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:05:06 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-lihc-tabular-open.TopBench
TopBench Dataset
TopBench is a benchmark for predictive reasoning over tabular data. Each example asks a model to infer an unobserved outcome, decision, treatment effect, or ranked/filtering result from historical tables and a natural-language query.
Layout
single_point_prediction/
decision_making/
treatment_effect_analysis/
ranking_and_filtering/
Each task directory contains dataset folders with history.csv, query JSON files, and metadata. The ranking_and_filtering task… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/TopBench.tcga-skcm-tabular-open
TCGA-SKCM — Tabular (Open Access)
Open-access TCGA-SKCM data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:17:45 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-skcm-tabular-open.bike-sharing-tabular
Bike Sharing Demand - Hourly (Poisson)
A ready-to-use copy of the UCI Bike Sharing Dataset (hourly granularity,
17,379 × 17), accompanied by baseline metrics from an 8-architecture tabular
modelling pipeline for direct comparison.
Originally collected and published by Fanaee-T & Gama (2014). Source:
UCI ML Repository id 275.
At a glance
Field
Value
Rows
17,379 hourly observations
Time range
Jan 2011 - Dec 2012
Columns
17 (16 features + 1 target)… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/bike-sharing-tabular.tcga-prad-tabular-open
TCGA-PRAD — Tabular (Open Access)
Open-access TCGA-PRAD data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:14:03 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-prad-tabular-open.tcga-blca-tabular-open
TCGA-BLCA — Tabular (Open Access)
Open-access TCGA-BLCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:45:21 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-blca-tabular-open.tcga-ov-tabular-open
TCGA-OV — Tabular (Open Access)
Open-access TCGA-OV data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:11:26 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-ov-tabular-open.tcga-brca-tabular-open
TCGA-BRCA — Tabular (Open Access)
Open-access TCGA-BRCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:47:01 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-brca-tabular-open.tcga-ucec-tabular-open
TCGA-UCEC — Tabular (Open Access)
Open-access TCGA-UCEC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:24:58 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-ucec-tabular-open.tabular-errors-v1
TabFix multilingual table error pairs — version 2.0
This release keeps 18 business error categories and separates executable deterministic detection from two residual neural categories: text.encoding and text.spelling. The same repository and family-disjoint splits are retained.
Split
Records
Open-vocabulary views
train
27948
3260
validation
17127
1844
test
32776
3540
The seven string columns remain id, split, family_id, clean_xml, corrupt_xml, errors… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/tabular-errors-v1.tabular-benchmark-797-classificationtcga-lgg-tabular-open
TCGA-LGG — Tabular (Open Access)
Open-access TCGA-LGG data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:02:58 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-lgg-tabular-open.tcga-sarc-tabular-open
TCGA-SARC — Tabular (Open Access)
Open-access TCGA-SARC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:16:39 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-sarc-tabular-open.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/tabular-benchmark.tcga-esca-tabular-open
TCGA-ESCA — Tabular (Open Access)
Open-access TCGA-ESCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:54:06 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-esca-tabular-open.tabular-logs-and-datasetstcga-tgct-tabular-open
TCGA-TGCT — Tabular (Open Access)
Open-access TCGA-TGCT data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:21:25 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-tgct-tabular-open.
