visual reasoning
proj_only_llava_deepseek_r1_distill_llama3_8b_reasoning_visual_cot_1000_samples_10_epochsproj_only_llava_deepseek_r1_distill_llama3_8b_reasoning_visual_cot_10000_samplesproj_only_llava_deepseek_r1_distill_llama3_8b_reasoning_visual_cot_1000_samplesvisual_reasoning_stage_1Visual-ReasoningIndustrial-Visual-Reasoning-GenAIvlm-mix-resume-visual-reasoning-expert-step100vlm-mix-visual-reasoning-expert-step100
Visual-Searchvisual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.Visual-JigsawVisualReasoningTracergrounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.spatial-visual-reasoning-66k
