Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes3.6k downloads2mo agoHugging Face02vector-institute /open-pmc-18m OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.image10M<n<100M7 likes2.9k downloads5mo agoHugging Face03vector-institute /sonic-o1 &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding 🎯 What is SONIC-O1? The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts. Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.audiovisual-question-answering1K<n<10K5 likes971 downloads5mo agoHugging Face04Maktabati /openiti-vectors OpenITI Vector Database — maktabati.ai 🇬🇧 English This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG). Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata. Statistics: 4,696,703 chunks 8,943 works (primary editions only, status=pri from OpenITI TSV) approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.tabular1M<n<10M0 likes949 downloads4mo agoHugging Face05Alfaxad /vector-100k VectorOS Vector 100k SimSat VLM Dataset VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M. The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.imagevisual-question-answering100K<n<1M1 likes520 downloads5mo agoHugging Face06vector-institute /open-pmc OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.image1M<n<10M9 likes519 downloads2y agoHugging Face07vectoryyyy /onevision1.5image100K<n<1M0 likes471 downloads8mo agoHugging Face08vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3text100K<n<1M0 likes432 downloads1y agoHugging Face09Maktabati /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes401 downloads4mo agoHugging Face10Delta-Vector /Tauri-RL-Plaintext-System-V2textn<1K1 likes393 downloads2mo agoHugging Face11vectorzhou /AIME_2024_DeepSeek_R1_0528_Temp_1.0_L_16384Responses of deepseek-ai/DeepSeek-R1-0528 for AIME 2024 (original dataset: Maxwell-Jia/AIME_2024). Generation temperature is set to 1.0 and maximum token is set to 16384. textn<1K0 likes370 downloads1y agoHugging Face12Delta-Vector /Tauri-RL-Stylestextn<1K0 likes366 downloads7mo agoHugging Face13Delta-Vector /Tauri-RL-Markdown-System-V2textn<1K0 likes363 downloads2mo agoHugging Face14DB-Edinburgh /VectorBenchmarkgatedtabular1M<n<10M3 likes351 downloads3d agoHugging Face15Delta-Vector /Tauri-Opus-Accepted-GPT-Rejected-Opus-Writing-Promptstext1K<n<10K3 likes256 downloads1y agoHugging Face16BinPrey /x86-instruction-test-vectors x86-64 Instruction Test Vectors Ground truth behavior of individual x86-64 instructions, captured by executing every encoding on real hardware and recording the resulting register and flag state. This is measured silicon behavior, not a model and not an emulator, so it also reflects implementation specific results such as the values an instruction leaves in flags that the architecture documents as undefined. How it was generated Each test case is produced by the… See the full description on the dataset page: https://huggingface.co/datasets/BinPrey/x86-instruction-test-vectors.text100M<n<1B1 likes242 downloads12d agoHugging Face17moss-vector-714 /GeoFidelity-Bench GeoFidelity-Bench GeoFidelity-Bench: Evaluating Block-Conditioned Geographic Fidelity in Text-to-Image Street-View Generation Kaizhen Tan, NeurIPS 2026. Contact: kt3275@nyu.edu. Version 3.1.0 aligns the public metadata with the original experiments: 112 named street blocks, 25 cities, 23 countries, 7,563 reference assignments covering 7,433 distinct Mapillary image IDs, and 16,128 generated images. The generations cover six models, six prompt conditions, and four samples per… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.imagetext-to-image10K<n<100K0 likes237 downloads15d agoHugging Face18dlxjj /NLPP_vectortextn<1K0 likes230 downloads1y agoHugging Face19ajaysri /pi07_cable_three_vector_v1 Three-holder cable routing with vector goals Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy. Property Value Source recordings 140 Subtask episodes 840 Frames 335,897 Sampling rate 100 Hz Robot ARX bimanual Recorded state / action 14 / 14 dimensions Video resolution 448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.tabular100K<n<1M0 likes206 downloads1mo agoHugging Face20ajaysri /route_red_yellow_vector_subtasks_pi05 Route Red-Yellow Vector Subtasks for pi0.5 This is a LeRobot v2.1 transformation of DistantSky/route at commit aced1e96f6b8bf98ffaa0754407636e82442b084. The dataset contains 212 real-robot episodes, 135699 frames, three cameras, and 14-dimensional actions at 100 Hz. Conditioning Only observation.images.video_overhead is modified. The left and right videos are byte-identical to the source dataset. The overhead image receives one fixed selected-connector pose glyph… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_red_yellow_vector_subtasks_pi05.tabularrobotics100K<n<1M0 likes175 downloads3mo agoHugging Face21vector-institute /hotpotqatext100K<n<1M0 likes172 downloads1y agoHugging Face22Ma-Vector /MobileGym-ConAct-Trajectories MobileGym-ConAct-Trajectories Dataset Viewer · MobileGym · MemGUI-Agent · Paper Abstract MobileGym-ConAct-Trajectories is a release of successful mobile GUI-agent rollouts collected in the MobileGym simulator. Each trajectory is selected from judge-verified rollouts using a deterministic per-task rule: retain the shortest structurally valid success, then break ties by source run and episode ID. The release preserves screenshots, the rendered prompt supplied to the… See the full description on the dataset page: https://huggingface.co/datasets/Ma-Vector/MobileGym-ConAct-Trajectories.imageimage-text-to-text1K<n<10K0 likes165 downloads3mo agoHugging Face23KShivendu /nanobeir-colbert-token-vectors NanoBEIR ColBERT token vectors Per-token ColBERT vectors for all 13 NanoBEIR datasets, from two late-interaction models: folder model revision encoder dtype lateon/ lightonai/LateOn 62911e105059585d244384c7d17826e35f669c17 fp32 iso/ topk-io/Iso-ModernColBERT b95d9608ad424fdd7bfd045001576f59cbb89a98 bf16 pplx-late-0.6b/ perplexity-ai/pplx-embed-v2-late-0.6b dd4e95b836a73f6f0c32e46ea127c0b86b02169e bf16 Each model covers 56,723 documents (7,585,402 kept… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/nanobeir-colbert-token-vectors.texttext-retrieval100K<n<1M0 likes161 downloads2d agoHugging Face24oitnews /rss_vectorstext100K<n<1M0 likes152 downloads1y agoHugging Face25derenrich /wiki-image-vectorsVector embeddings of images on Wikipedia. Images were embedded with "google/siglip2-base-patch16-384" The "002" partition contains all of the quality images. tabular1M<n<10M0 likes150 downloads9d agoHugging Face26oitnews /googlenews_vectorstext1K<n<10K0 likes145 downloads1y agoHugging Face27Kandil7 /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes145 downloads4mo agoHugging Face28oitnews /newsdataio_vectorstext10K<n<100K0 likes140 downloads1y agoHugging Face29vectorzhou /AIME_2024_Qwen3_235B_A22B_Temp_1.0_L_16384Responses of Qwen/Qwen3-235B-A22B for AIME 2024 (original dataset: Maxwell-Jia/AIME_2024). Generation temperature is set to 1.0 and maximum token is set to 16384. textn<1K1 likes140 downloads1y agoHugging Face30mikronai /VectorEdits VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics NOTE: Currently only test set has generated labels, other sets will have them soon Find the details in our paper: VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics Github repository: JosefKuchar/vector-edits We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language… See the full description on the dataset page: https://huggingface.co/datasets/mikronai/VectorEdits.text100K<n<1M7 likes122 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.