datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabularMath
📊 TabularMath
TabularMath is a tabular mathematical reasoning benchmark introduced in TabularMath: Understanding Math Reasoning over Tables with Large Language Models. It is built via AUTOT2T, a neuro-symbolic pipeline that automatically transforms math word problems into verified tabular reasoning tasks, enabling scalable evaluation without manual table annotation.
TabularMath jointly assesses reasoning accuracy, information retrieval over complex table structures, and… See the full description on the dataset page: https://huggingface.co/datasets/kevin715/TabularMath.Artistic_Landscape
Artistic Landscape Dataset
Welcome to the Artistic Landscape.
Overview
Artistic Landscape is a curated collection of synthetically generated imagery designed to explore conceptual combinations across a wide variety of art styles, media, and visual characteristics. It is intended to be exploratory in scope, comparative across many aesthetic dimensions, and practical for downstream research workflows where a large, structured visual vocabulary is useful.
In total, we used… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/Artistic_Landscape.yash-gym-tabular-dataset
Yash Gym Tabular Dataset
Dataset Summary
This dataset contains information on 30 unique gym machines with 5 consistent features and a binary target (Upper/Lower).It includes:
original: 30 manually collected samples
augmented: ~300 synthetic samples created with jitter, SMOTE-NC, MixUp, and CTGAN.
Intended Use
Educational dataset for tabular ML tasks, demonstrating preprocessing + augmentation.Not suitable for prescribing exercise or medical advice.… See the full description on the dataset page: https://huggingface.co/datasets/ysakhale/yash-gym-tabular-dataset.yalta_ai_tabular_dataset
YALTAi Tabular Dataset
353 page images of historical documents with tabular layouts, mostly notarial registers, with bounding-box annotations for four zone types: Col, Header, Marginal and text. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem rather than a pixel… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_tabular_dataset.weasis-tabular-benchmark
Weasis Medical Imaging GUI Benchmark (Tabular Format)
Dataset Description
This dataset contains 267 end-to-end GUI automation tasks for the Weasis medical imaging viewer in tabular format, where each row represents one complete task with all associated data.
Dataset Summary
Total Tasks: 267
Total Images: 202
Format: Tabular (each row = one task)
Application: Weasis Medical Imaging Viewer
Resolution: 1920x1080
Data Structure
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/rishuKumar404/weasis-tabular-benchmark.test_composite_and_tabular_datasetMMTU
Benchmark Datasets used for "Compositional Condition Question Answering in Tabular Understanding"
This repository contains the benchmark used in "Compositional Condition Question Answering in Tabular Understanding". Specifically, we provide the datasets in two zip files: data.zip and picked.zip.
data.zip:
This zip file contains four JSON files, including IE.json, CR.json, CC.json and RC.json, each containing all the samples of one category of QA pairs. Each contains following… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/MMTU.train_composite_and_tabular_datasetL1_tabular_data
Dataset Card for "L1_tabular_data"
More Information needed
