datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsm8k-rendered-vlm
Rendered GSM8K-VL Dataset
Rendered GSM8K-VL is a multimodal math-reasoning dataset for vision-language model evaluation.Each example links:
a GSM8K word problem (question)
the final numeric answer (answer)
cleaned chain-of-thought style reasoning (reasoning)
a rendered image path (image)
This dataset is intended for controlled experiments comparing text-only and image-based reasoning behavior.
Canonical Dataset Artifact
The official dataset release uses:… See the full description on the dataset page: https://huggingface.co/datasets/RodelaG/gsm8k-rendered-vlm.coding-model-rendered-qa
Rendered QA Dataset: Code & Text (700K)
Instruction-tuning dataset with optional rendered images for vision-language models.
Sources
Source
Samples
Has Context Image
OpenCoder Stage 2
436K
educational_instruct only
InstructCoder
108K
Yes (code input)
OpenOrca
200K
No (text-only)
Schema
Column
Type
Description
prompt
string
Instruction/question
prompt_image
Image?
Rendered prompt (optional)
context
string?
Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.
