Team Ai
20 results

multi-modal

SWE-bench /SWE-bench_Multimodal SWE-bench Multimodal Dataset Summary SWE-bench Multimodal is a dataset that tests systems' ability to resolve real-world GitHub issues in visual software domains. Unlike the original SWE-bench, which is Python-only and text-only, every task instance here comes from a JavaScript or TypeScript repository and carries at least one image asset — a screenshot, a screen recording, a diagram, or a rendering of incorrect output. The dataset collects 612 Issue-Pull Request pairs from 17… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Multimodal.textn<1K13 likes15k downloads2mo agoHugging Face0001AMA /multimodal_data_annotator_datasetMaterials dataset consisting of spatial and time resolved versions of the same object. Specially curated for the annotator such that for each object, time resolved signal may be viewed alongside the RGB and for different graphs/forms image10K<n<100K0 likes12k downloads8mo agoHugging Facemultimodal-reasoning-lab /Zebra-CoT Zebra‑CoT A diverse large-scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games. Dataset Structure Each example in Zebra‑CoT consists of: Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.imageany-to-any100K<n<1M80 likes11k downloads8mo agoHugging Facealibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes10k downloads3mo agoHugging Faceosunlp /Multimodal-Mind2Web Dataset Summary Multimodal-Mind2Web is the multimodal version of Mind2Web, a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. In this dataset, we align each HTML document in the dataset with its corresponding webpage screenshot image from the Mind2Web raw dump. This multimodal version addresses the inconvenience of loading images from the ~300GB Mind2Web Raw Dump. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web.image10K<n<100K99 likes9.6k downloads2y agoHugging Faceomegalabsinc /omega-multimodal OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research Introduction The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation. With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.tabularvideo-text-to-text60 likes9k downloads1y agoHugging Face