Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShareLab-SII /thinking_droid_lerobot_output_qwen3vlimage1M<n<10M0 likes3.4k downloads6mo agoHugging Face02OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes2.5k downloads8mo agoHugging Face03krishnateja95 /ImageNet-Think ImageNet-Think 250K ImageNet-Think 250K is a large-scale synthetic multimodal reasoning dataset containing of 250,000 images sampled from ImageNet-21K dataset. For each image, we provide a prompt and two different step-by-step reasoning tokens and outputs (answers), enabling evaluation and training for Vision Language Models on reasoning tasks. This dataset is primarily designed for research on multimodal summarization. Installation & Setup Before downloading… See the full description on the dataset page: https://huggingface.co/datasets/krishnateja95/ImageNet-Think.imagesummarization100K<n<1M4 likes2.2k downloads1y agoHugging Face04ShareLab-SII /thinking_furniture_bench_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes2.1k downloads6mo agoHugging Face05NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.5k downloads8mo agoHugging Face06OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.5k downloads7mo agoHugging Face07ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face08RuoliuYang /textlatent_zebra_thinkmorph_armAB Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph 35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset, question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json). armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3). armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17). messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.image100K<n<1M0 likes939 downloads3mo agoHugging Face09Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes877 downloads8mo agoHugging Face10UCSC-VLAA /VLAA-Thinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄 Arxiv • 💻 Code 🤗 VLAA-Thinker Family • 🤔 VLAA-Thinking Dataset 🤗 VLAA-Thinker-Qwen2.5-3B • 🤗 VLAA-Thinker-Qwen2.5-7B Both VLAA-Thinker-Qwen2.5-3B and VLAA-Thinker-Qwen2.5-7Bachieve SOTA performance on OpenCompass Multimodal Reasoning Leaderboard as of April 7th, 2025. Contents Quick Start 🚀… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLAA-Thinking.documentvisual-question-answeringn<1K20 likes747 downloads1y agoHugging Face11ShareLab-SII /thinking_fmb_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes686 downloads7mo agoHugging Face12MBZUAI /ThinkGeo ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks imagen<1K4 likes661 downloads6mo agoHugging Face13AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes464 downloads3mo agoHugging Face14yrlyrl /spatial-mmcot-thinkmorph_spatial_nav Spatial MMCoT v1 · thinkmorph_spatial_nav ThinkMorph (arXiv:2510.27492) Spatial_Navigation: FrozenLake navigation on 3x3 to 6x6 grids. The one input image is the grid, and the question text is the same on every row (upstream's wording, unchanged: it asks for the moves in \boxed{{}}, an unformatted template left in upstream; the read-back boxes the moves while <answer> holds them bare), so the maze exists only in the image. The first thought describes the grid (start, goal… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-thinkmorph_spatial_nav.imagevisual-question-answering10K<n<100K0 likes402 downloads7d agoHugging Face15dans25275 /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/dans25275/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes401 downloads8mo agoHugging Face16yrlyrl /spatial-mmcot-thinkmorph_jigsaw Spatial MMCoT v1 · thinkmorph_jigsaw ThinkMorph (arXiv:2510.27492) Jigsaw_Assembly: a picture cut into numbered parts, shown in an order that may be shuffled; decide the arrangement, draw the assembled picture, read it back, answer. According to the ThinkMorph paper (arXiv:2510.27492, appendix on data generation), GPT-4.1 wrote the plan and the read-back from the question and the ground-truth answer, and was told not to reveal the answer. The pictures come from 3 corpora: 3,201… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-thinkmorph_jigsaw.imagevisual-question-answering10K<n<100K0 likes401 downloads7d agoHugging Face17OpenDataArena /MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking MMFineReason-SFT-586K The Hardest 33% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1). Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning. 🎯 Key Highlights 586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.image100K<n<1M6 likes375 downloads8mo agoHugging Face18olob0 /finevision-mini-thinking FineVision-mini Thinking FineVision-mini is a slice I made of HuggingFaceM4/FineVision: 169 of its image subsets, 101,321 rows (~40 GB) out of FineVision's 24.2M rows / 4.65 TB (about 0.4% of the rows, 0.9% of the bytes), sampled with a fixed seed. The 16 text-only subsets were left out. This dataset is that slice, fully translated and augmented with reasoning, published in increments: each batch processes more rows of FineVision-mini and is appended here, until the whole slice… See the full description on the dataset page: https://huggingface.co/datasets/olob0/finevision-mini-thinking.imagevisual-question-answering10K<n<100K0 likes345 downloads24d agoHugging Face19OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes243 downloads8mo agoHugging Face20ThinkMorph /Spatial_Navigation 🌟 This repo contains part of the training dataset for model ThinkMorph-7B. Dataset Description We create an enriched interleaved dataset centered on four representative tasks requiring varying degrees of visual engagement and cross-modal interactions, including Jigsaw Assembly, Spatial Navigation, Visual Search and Chart Refocus. Statistics Dataset Usage Data Downloading… See the full description on the dataset page: https://huggingface.co/datasets/ThinkMorph/Spatial_Navigation.imageany-to-any1K<n<10K4 likes240 downloads11mo agoHugging Face21ThinkMorph /Jigsaw_Assembly 🌟 This repo contains part of the training dataset for model ThinkMorph-7B. Dataset Description We create an enriched interleaved dataset centered on four representative tasks requiring varying degrees of visual engagement and cross-modal interactions, including Jigsaw Assembly, Spatial Navigation, Visual Search and Chart Refocus. Statistics Dataset Usage Data Downloading… See the full description on the dataset page: https://huggingface.co/datasets/ThinkMorph/Jigsaw_Assembly.imageany-to-any1K<n<10K0 likes194 downloads11mo agoHugging Face22AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes185 downloads3mo agoHugging Face23VIBE-Benchmark /BAGEL-thinkimage1K<n<10K0 likes175 downloads8mo agoHugging Face24UoM-CS-NeuroSymbolicAI /pred_qwen3vl_think_10kimage1K<n<10K0 likes172 downloads6mo agoHugging Face25purefall /shotpath-qwen3vl-thinking-cot-20260714image1K<n<10K0 likes168 downloads3mo agoHugging Face26appletea2333 /ThinkEditimage0 likes147 downloads9mo agoHugging Face27AmirhoseinGH /mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes138 downloads3mo agoHugging Face28AmirhoseinGH /mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes131 downloads3mo agoHugging Face29AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes130 downloads3mo agoHugging Face30ThinkMorph /Visual_Search 🌟 This repo contains part of the training dataset for model ThinkMorph-7B. Dataset Description We create an enriched interleaved dataset centered on four representative tasks requiring varying degrees of visual engagement and cross-modal interactions, including Jigsaw Assembly, Spatial Navigation, Visual Search and Chart Refocus. Statistics Dataset Usage Data Downloading… See the full description on the dataset page: https://huggingface.co/datasets/ThinkMorph/Visual_Search.imageany-to-any1K<n<10K1 likes117 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.