Team Ai
20 results

INSIGHT

NTT-hil-insight /SlideVQAgated SlideVQA SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images 📖 arXiv 🌐 github We introduce a new document VQA dataset, SlideVQA, for tasks wherein given a slide deck composed of multiple slide images and a corresponding question, a system selects a set of evidence images and answers the question. Citation and contact If you use this dataset, please cite our work: @inproceedings{SlideVQA2023, author = {Ryota Tanaka and… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/SlideVQA.imagevisual-question-answering10K<n<100K21 likes2.3k downloads2y agoHugging Faceinsight /locomoaudio100K<n<1M1 likes1.9k downloads7mo agoHugging Faceinsight /cache0 likes1.7k downloads7mo agoHugging Facem-Just /InSight-doc-SFT-18k InSight-doc-SFT-18k Agentic Visual Perception for Long-Document Understanding 📄 Paper | 💻 Code | 🤗 Model | 🎯 RL Data | 🎬 Replay Demo | 🚀 Live Demo Understand the big picture.&nbsp; Focus on the right details.&nbsp; Answer from the evidence. InSight-doc-SFT-18k is the supervised fine-tuning corpus used to train the InSight-doc long-document understanding agent. Each example is a complete multimodal trajectory: the agent starts from low-resolution document… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-SFT-18k.imagevisual-question-answering10K<n<100K3 likes1.4k downloads2mo agoHugging Faceparagon7060 /InsightBench-Assets0 likes1k downloads4mo agoHugging Facelaion /empathic-insights-voice1 likes985 downloads1y agoHugging Face