Team Ai
20 results

Vision-language

mlfu7 /Touch-Vision-Language-Dataset A Touch, Vision, and Language Dataset for Multimodal Alignment by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution). [Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation] This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/Touch-Vision-Language-Dataset.11 likes956 downloads1y agoHugging Facemicrosoft /VISION_LANGUAGEA key question for understanding multimodal vs. language capabilities of models is what is the relative strength of the spatial reasoning and understanding in each modality, as spatial understanding is expected to be a strength for multimodality? To test this we created a procedurally generatable, synthetic dataset to testing spatial reasoning, navigation, and counting. These datasets are challenging and by being procedurally generated new versions can easily be created to combat the effects… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VISION_LANGUAGE.image10K<n<100K7 likes534 downloads2y agoHugging FaceNOVA-vision-language /calame-pt CALAME-PT Context-Aware LAnguage Modeling Evaluation for Portuguese CALAME-PT is a PT benchmark composed of small texts (contexts) and their respective last words. These contexts should, in theory, contain enough information so that a human or a model is capable of guessing its last word - without being too specific and/or too ambiguous. Composition CALAME-PT is composed of 2 "sets" of data - handwritten and generated. Handwritten Set: contains 406… See the full description on the dataset page: https://huggingface.co/datasets/NOVA-vision-language/calame-pt.text1K<n<10K4 likes242 downloads3y agoHugging Facebeatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging Facehassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K7 likes148 downloads7mo agoHugging Faceipranavks /visionlanguagemodelogimagen<1K8 likes134 downloads1y agoHugging Face