deepvision
DRR_dataDeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/Devilishcode/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/JamesGoGo/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/blsmash044/DeepVision-103K.deepvision-vlm-predictions
VLM Evaluation Predictions — Zeroshot-DeepVision-24k
Model evaluation predictions and metrics from the Bengali Math VQA pipeline.
Folder structure
Baseline evaluation protocol
Base model output is constrained via a system-level instruction to output only the
Bengali MCQ option letter (ক/খ/গ/ঘ), ensuring format parity with the fine-tuned model.
