datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRR_dataDeepVision-103K
š DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /⦠See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/DeepVision-103K.DeepVision-103K
š DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /⦠See the full description on the dataset page: https://huggingface.co/datasets/Devilishcode/DeepVision-103K.DeepVision-103K
š DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /⦠See the full description on the dataset page: https://huggingface.co/datasets/JamesGoGo/DeepVision-103K.DeepVision-103K
š DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /⦠See the full description on the dataset page: https://huggingface.co/datasets/blsmash044/DeepVision-103K.deepvision-vlm-predictions
VLM Evaluation Predictions ā Zeroshot-DeepVision-24k
Model evaluation predictions and metrics from the Bengali Math VQA pipeline.
Folder structure
Baseline evaluation protocol
Base model output is constrained via a system-level instruction to output only the
Bengali MCQ option letter (ą¦/ą¦/ą¦/ą¦), ensuring format parity with the fine-tuned model.
deepvisiontools-demo-datasets
Description
Those are 3 datasets for deepvisiontools library demo. deepvisiontools homepage : https://forge.inrae.fr/ue-apc/librairies/python/deepvisiontools
Datasets
The dice dataset was downloaded from Kaggle : https://www.kaggle.com/datasets/nellbyler/d6-dice
The VegannSubDataset is a small portion from : https://zenodo.org/records/7636408
The coco_6cls_subset was obtained from : https://universe.roboflow.com/nan-ixwz3/coco-y1tdb
deepvision-zero-shot-20kdeepvision_datasetsdeepvisiontools_tutorialsDeepVision_ZS_Predictions
VLM Evaluation Predictions
Prediction outputs from the Zeroshot-DeepVision-24k Bengali Math VQA evaluation pipeline.
Structure
{model_short}/
baseline-testing/ ā predictions BEFORE fine-tuning
post-finetune-testing/ ā predictions AFTER fine-tuning
Generated by the reusable Kaggle evaluation notebook.
deepvision-atlas-data
