Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01omegalabsinc /omega-multimodal OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research Introduction The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation. With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.tabularvideo-text-to-text60 likes9.2k downloads1y agoHugging Face02Perle-ai /multimodal-ct-radiology-reports Perle AI Multi-phase CECT and CT with Radiology Reports Summary A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering. The release has three configurations: Config Modality Subjects Pairing cect_3phase 3-phase contrast-enhanced abdominal CT (DICOM) 5 per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.tabularimage-classificationn<1K5 likes8.9k downloads5mo agoHugging Face03MultimodalUniverse /plasticc--- description: 'The Photometric LSST Astronomical Time-Series Classification Challenge (PLAsTiCC) is a community-wide challenge to spur development of algorithms to classify astronomical transients. The Large Synoptic Survey Telescope (LSST) will discover tens of thousands of transient phenomena every single night. To deal with this massive onset of data, automated algorithms to classify and sort astronomical transients are crucial. ' homepage: https://zenodo.org/records/2539456… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/plasticc.tabular1K<n<10K1 likes7.5k downloads2y agoHugging Face04BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face05lerobot /stanford_kuka_multimodal_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 3000, "total_frames": 149985, "total_tasks": 1, "total_videos": 3000, "total_chunks": 3, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:3000" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.tabularrobotics100K<n<1M2 likes3.3k downloads1y agoHugging Face06electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.6k downloads2mo agoHugging Face07FineEnvs /multilingual-multimodal-rl-runs Multilingual Multimodal RL runs The evidence behind the two Kannada GRPO runs in the Multilingual Multimodal Envs collection: what the model wrote for every held-out item at every checkpoint, the curves, the training logs, and the exact scripts that ran. If a number on one of the model cards looks wrong, this is where to check it. Run Model Trained on Held-out set Base → selected checkpoint asr-kannada gemma-4-E4B-it-kannada-asr-grpo all 2,282 FLEURS kn_in train… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/multilingual-multimodal-rl-runs.tabular10K<n<100K2 likes1.8k downloads4d agoHugging Face08ExylosAi /egocentric-vr-capture-20h-multimodal-sample Egocentric VR Capture — 20-Hour Multimodal Inspection Sample 195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.tabularrobotics1M<n<10M0 likes1.6k downloads18d agoHugging Face09Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes751 downloads4mo agoHugging Face10Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K2 likes690 downloads1mo agoHugging Face11ai-team-core /adapter-based-multimodal-fusion Falcon-Audio Training Dataset Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes. tabular100K<n<1M0 likes590 downloads5mo agoHugging Face12BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes477 downloads2y agoHugging Face13pakkinlau /multi-modal-derived-brain-network PPMI Connectivity Graphs — HF Staging (Derivatives) This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include: ROI time-series (*_desc-timeseries_parc-<name>.mat) Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat) JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags) Contents data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.tabularn<1K0 likes441 downloads1y agoHugging Face14MultimodalUniverse /gaia--- description: 'Spectral (BP/RP), photometric, and astrometric dataset based on Gaia DR3. ' homepage: https://www.cosmos.esa.int/web/gaia/dr3 version: 1.0.0 citation: "% % ACKNOWLEDGEMENTS\n% If you have used Gaia DR3 data in your research, \ please use the following acknowledgement:\n% \n% This work has made use of data \ from the European Space Agency (ESA) mission\n% {\it Gaia} (\url{https://www.cosmos.esa.int/gaia}),\ \ processed by the {\it Gaia}\n% Data Processing and Analysis… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/gaia.tabular100K<n<1M1 likes375 downloads2y agoHugging Face15MultimodalUniverse /hsc--- description: 'Image dataset based on HSC SSP PRD3. ' homepage: https://hsc-release.mtk.nao.ac.jp/doc/ version: 1.0.0 citation: "% CITATION\n@article{Aihara_2017,\n title={The Hyper Suprime-Cam SSP \ Survey: Overview and survey design},\n volume={70},\n ISSN={2053-051X},\n \ url={http://dx.doi.org/10.1093/pasj/psx066},\n DOI={10.1093/pasj/psx066},\n \ number={SP1},\n journal={Publications of the Astronomical Society of Japan},\n \ publisher={Oxford University Press… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/hsc.tabular10K<n<100K0 likes255 downloads2y agoHugging Face16phi-9 /ego-multimodal ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR) Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited. A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.tabularrobotics10K<n<100K2 likes248 downloads7mo agoHugging Face17trucyberlab /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M1 likes231 downloads3mo agoHugging Face18June30916 /multimodality-poc-llama31-ruler16k Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K Raw pre-RoPE query and hidden-state tensors captured during prefill, used to study whether the per-(layer, kv_head) query distribution is unimodal Gaussian (the assumption underpinning Expected Attention's MGF closed-form in kvpress). What's in here 65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5 prompts). Each file (~414 MB) contains: field dtype shape meaning hidden float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.tabularfeature-extractionn<1K0 likes227 downloads5mo agoHugging Face19hashmortar /multimodal-annual-reports Multimodal Annual Reports A document question-answering benchmark built from 20 complete corporate annual and integrated reports. It contains 595 English questions with reference answers, source evidence, original PDFs, reviewed HTML and Markdown representations, and figure/table crops. Questions require interpreting narratives, tables, and non-tabular visuals, including Japanese and French sources. MIT covers original benchmark contributions only. Source reports and their… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/multimodal-annual-reports.imagedocument-question-answering1K<n<10K0 likes195 downloads4d agoHugging Face20Multimodal-Fatima /VQAv2_test_no_image Dataset Card for "VQAv2_test_no_image" More Information needed tabular100K<n<1M0 likes185 downloads3y agoHugging Face21MultimodalUniverse /sdss--- description: 'Spectra dataset based on SDSS-IV. ' homepage: https://www.sdss.org/ version: 1.0.0 citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\ % \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.tabular100K<n<1M3 likes185 downloads2y agoHugging Face22superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face23beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging Face24hainh22 /multi-modal-peg-in-square-hole-testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5", "total_episodes": 51, "total_frames": 8074, "total_tasks": 1, "total_videos": 153, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:51" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-test.tabularrobotics1K<n<10K0 likes166 downloads2y agoHugging Face25MultimodalUniverse /jwst--- description: 'Image dataset based on a combination of JWST deep fields from DJA: CEERS, NGDEEP, JADES, PRIMER ' homepage: https://dawn-cph.github.io/dja/index.html version: 1.1.0 citation: "% % ACKNOWLEDGEMENTS\n% % From: https://dawn-cph.github.io/dja/index.html\n\ % We kindly request all scientific papers based on data or products downloaded from \ the Dawn JWST Archive (DJA) to include the following acknowledgement:\n% \n% (Some \ of) The data products presented herein were… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/jwst.tabular10K<n<100K4 likes156 downloads2y agoHugging Face26introvoyz042 /egocentric-vr-capture-1h-multimodal-sample Egocentric VR Capture — 1-Hour Multimodal Inspection Sample 13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.tabularrobotics100K<n<1M0 likes148 downloads1mo agoHugging Face27fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes148 downloads22d agoHugging Face28hainh22 /multi-modal-peg-in-square-holeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5", "total_episodes": 201, "total_frames": 28616, "total_tasks": 1, "total_videos": 603, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:201" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole.tabularrobotics10K<n<100K0 likes135 downloads2y agoHugging Face29neurips26 /MultimodalUnlearningEvalBenchmark 🧠 Multimodal Unlearning Evaluation Benchmark 📌 Overview This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning. It supports reproducibility of results in: Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026) 📊 Contents File Description 📄 multimodal_results.json Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench) 📄 unimodal_results.json CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.tabulartext-generationn<1K1 likes135 downloads5mo agoHugging Face30fluid-concepts /multimodal-expert-instruction-samplesgated Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside. ▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.audion<1K1 likes134 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.