datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChartQA
Dataset Card for "ChartQA"
More Information needed
ChartQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of ChartQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{masry2022chartqa,
title={ChartQA: A benchmark for question answering about charts with visual and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ChartQA.ChartQAIf you wanna use the dataset, you need to download the zip file manually from the "Files and versions" tab.
Please note that this dataset can not be directly loaded with the load_dataset function from the datasets library.
If you want a version of the dataset that can be loaded with the load_dataset function, you can use this one: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images
But it doesn't contain the chart images. Hence, you will still need to use the images stored in… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQA.ChartQAPro
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
🤗Dataset | 🖥️Code | 📄Paper
The abstract of the paper states that:
Charts are ubiquitous, as people often use them to analyze data, answer questions, and discover critical insights. However, performing complex analytical tasks with charts requires significant perceptual and cognitive effort. Chart Question Answering (CQA) systems automate this process by enabling models to interpret and reason with… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQAPro.chartqa_without_images
Dataset Card for "chartqa_without_images"
If you wanna load the dataset, you can run the following code:
from datasets import load_dataset
data = load_dataset('ahmed-masry/chartqa_without_images')
The dataset has the following structure:
DatasetDict({
train: Dataset({
features: ['imgname', 'query', 'label', 'type'],
num_rows: 28299
})
val: Dataset({
features: ['imgname', 'query', 'label', 'type'],
num_rows: 1920
})
test:… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images.VisRAG-Ret-Test-ChartQA
Dataset Description
This is a VQA dataset based on Charts from ChartQA dataset from ChartQA.
Load the dataset
from datasets import load_dataset
import csv
def load_beir_qrels(qrels_file):
qrels = {}
with open(qrels_file) as f:
tsvreader = csv.DictReader(f, delimiter="\t")
for row in tsvreader:
qid = row["query-id"]
pid = row["corpus-id"]
rel = int(row["score"])
if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-ChartQA.chartqagridline-chartqa
Adaption Charts P2 — Gold Chart-QA Dataset
A verified, quality-first chart question-answering dataset built for the
Adaption Labs AutoScientist Challenge (Part 2, Data Visualization track).
Two sources: a programmatically generated synthetic core
(correct-by-construction) and a hand-authored hardset built from real
public dashboards and reports.
At a glance
1415 rows total — 1317 synthetic + 98 hardset
7 chart types — bar, line, grouped_bar, stacked_bar, pie… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/gridline-chartqa.scientific-chart-qa-17k
Scientific Chart QA, 17,070 rows
A multimodal chart-interpretation dataset built around one idea: teaching a model when not to
answer matters as much as teaching it to answer.
One in seven questions here cannot be answered from its figure, and the correct response is
cannot be determined. Baseline vision-language models overwhelmingly guess a plausible-looking
number instead. That is the behaviour this set targets.
The four things worth… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/scientific-chart-qa-17k.ChartQA_small_preprocessedChartQAVLLM_ChartQAchartqapro_disco
ChartQAPro Mini Dataset
A stratified 494-sample subset of the ChartQAPro dataset for chart question answering evaluation. This mini version maintains the diversity of the full dataset while being suitable for quick benchmarking and testing.
Dataset Description
ChartQAPro_mini contains question-answer pairs from diverse chart types with balanced representation across:
Question Types: Factoid (55.9%), Conversational (16%), Fact Checking (12.8%), Multi Choice… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/chartqapro_disco.arabic_chartqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_chartqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_chartqa_ar_beir.ChartQA_beirThis is a copy of https://huggingface.co/datasets/jinaai/ChartQA reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/ChartQA_beir.ChartQADataset is converted from https://github.com/vis-nlp/ChartQA
vin là tập đã dịch các qa3000 là tập các chart đã dịch
Disclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use.
PlotQa_ChartQa_cleanVLLM_ChartQA_splitChartQA_Benetech_PlotQa_DVQA_combined_matcha_completechartqa-grpoChart-Sum-QAchartqa-dataset-statistachartqa-derender-3ChartQAProChartQADatasetV2ChartQA dataset demoBToks-visrag_indomain_ChartQA
BToks VisRAG ChartQA
This dataset repository contains Lance-format converted data used by the open-source reproduction code for Bottleneck Tokens for Unified Multimodal Retrieval (arXiv:2604.11095).
Source
Converted from openbmb/VisRAG-Ret-Train-In-domain-data.
Subset/view: ChartQA. This repository does not change upstream ownership, licensing, citation requirements, or usage restrictions.
Format
The data is stored as Lance tables for the… See the full description on the dataset page: https://huggingface.co/datasets/siyrus/BToks-visrag_indomain_ChartQA.ChartQAProVQA-lmms-lab-ChartQA-clean
Description
French translation of the lmms-lab/ChartQA dataset that we processed.
Citation
@article{masry2022chartqa,
title={ChartQA: A benchmark for question answering about charts with visual and logical reasoning},
author={Masry, Ahmed and Long, Do Xuan and Tan, Jia Qing and Joty, Shafiq and Hoque, Enamul},
journal={arXiv preprint arXiv:2203.10244},
year={2022}
}
adaption-dataviz-chartqa-original-11k
Chart QA with Misleading Charts
Questions about described line and bar charts, including charts with deliberately distorted axes, with answers that read the underlying values.
Rows
11,720
Domain
data visualization
Format
data.parquet, one row per example
Licence
cc-by-4.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-dataviz-chartqa-original-11k.realworld-chartqa
Dataset Card for RealWorld-ChartQA
Summary
RealWorld-ChartQA is a benchmark dataset for chart question answering (CQA), derived from real-world analytical narratives. It contains 205 manually validated multiple-choice question–answer pairs grounded in student-authored literate visualization notebooks. Unlike previous CQA datasets, RealWorld-ChartQA includes multi-view and interactive charts, along with questions rooted in ecologically valid analytical workflows.… See the full description on the dataset page: https://huggingface.co/datasets/maevehutch/realworld-chartqa.
