Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes9.4k downloads23h agoHugging Face02BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes8k downloads1d agoHugging Face03BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes5.9k downloads23h agoHugging Face04PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K49 likes5.2k downloads2y agoHugging Face05nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes4.8k downloads2y agoHugging Face06Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face07PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face08Rapidata /human-alignment-preferences-images Rapidata Image Generation Alignment Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.imagetext-to-image10K<n<100K17 likes658 downloads2y agoHugging Face09PKU-Alignment /BeaverTails-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. 1. Usage If you want to use load_dataset(), you can directly use as follows: from datasets import load_dataset train_dataset = load_dataset('PKU-Alignment/BeaverTails-V', name='animal_abuse')['train'] eval_dataset = load_dataset('PKU-Alignment/BeaverTails-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-V.image10K<n<100K3 likes566 downloads2y agoHugging Face10yyshi0619 /ECOT-Alignment-900-Episodes ECOT Alignment 900 Episodes This dataset contains 900 rollout episodes generated by the original MiniVLA policy across all 90 LIBERO-90 tasks (10 distinct initial configurations per task). It was collected for supervised policy/reasoning alignment experiments. Successful and failed episodes are both included. Splits and counts Split Episodes Policy queries Training 720 13,093 Validation 180 3,229 Total 900 16,322 The split is task-stratified:… See the full description on the dataset page: https://huggingface.co/datasets/yyshi0619/ECOT-Alignment-900-Episodes.imagerobotics10K<n<100K0 likes492 downloads28d agoHugging Face11Rapidata /Flux_SD3_MJ_Dalle_Human_Alignment_Dataset NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset Rapidata Image Generation Alignment Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.imagetext-to-image10K<n<100K16 likes447 downloads2y agoHugging Face12macpaw-research /asset-alignment-pairs-905k Asset Alignment Pairs 905k Dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction". Large-scale dataset for rigid 3D asset alignment: given an independently generated 3D asset (src) and a target object (tgt), predict the rigid transformation that places the asset onto the target object. Each row is one source–target pair, rendered from three canonical orthogonal viewpoints with RGB, metric depth, camera extrinsics, and the ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-pairs-905k.image100K<n<1M0 likes376 downloads2mo agoHugging Face13PKU-Alignment /Align-Anything-TI2T-Instruction-100K Dataset Card for Align-Anything : Text-Image-to-Text Instruction-Following Subset Text+Image → Text Instruction-Following Dataset [🏠 Homepage] [🤗 Align-Anything Datasets] [🦫 Beaver-Vision-11B] Highlights Input & Output Modalities: Input: Text + Image; Output: Text 100K QA Pairs: Through refined construction based on constitutions, we obtained 103,012 QA pairs, with answers generated by GPT-4o. Beaver-Vision-11B: Leveraging our high-quality TI2T… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-TI2T-Instruction-100K.image100K<n<1M1 likes362 downloads2y agoHugging Face14macpaw-research /asset-alignment-reference-views Asset Alignment Reference Views Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction". Multi-view renderings of correctly assembled source–target pairs: each row shows one asset already aligned onto its target object, rendered from 12 orbiting viewpoints with RGB and depth. Where asset-alignment-pairs-905k shows the asset misaligned and supplies the transformation that fixes it, this dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.image10K<n<100K0 likes283 downloads2mo agoHugging Face15PKU-Alignment /PKU-SafeRLHF-VWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. 1. Usage If you want to use load_dataset(), you can directly use as follows: from datasets import load_dataset train_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V', name='animal_abuse')['train'] eval_dataset = load_dataset('PKU-Alignment/PKU-SafeRLHF-V'… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-V.image10K<n<100K6 likes282 downloads2y agoHugging Face16mtybilly /PubMedVision-Alignment-VQA PubMedVision-Alignment-VQA (flat single-image) Re-export of the PubMedVision_Alignment_VQA subset from FreedomIntelligence/PubMedVision processed for easier downstream consumption. Transformations vs. upstream Single-image rows only: rows with multiple images dropped (~22% of original) 9 rows with missing image files (upstream packaging gap; e.g. pmc_9_0.jpg is referenced but absent from images_*.zip) are also dropped conversations expanded into separate question and… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/PubMedVision-Alignment-VQA.imagevisual-question-answering100K<n<1M1 likes240 downloads5mo agoHugging Face17BDRC /ALL-BDRC-alignments Tibetan OCR — ALL-BDRC-alignments 79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come from the Asian Classics Input Project (ACIP) Sungbum corpus via the Asian Legacy Library (ALL), normalized to Unicode and manually matched to BDRC scans. This is the largest clean uchen woodblock set in the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments.imageimage-to-text10K<n<100K1 likes220 downloads1mo agoHugging Face18ivezakis /llava_med_alignment_500k_chunk_2image10K<n<100K0 likes193 downloads10mo agoHugging Face19visv-Bro /brain-ai-alignment-reproduction Brain–AI alignment reproduction artifacts Full-scale statistical audit of “Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories.” Paper: https://arxiv.org/abs/2507.01966 Audited source commit: https://github.com/FloyedShen/BrainAlign/tree/f3e8c78ed3ee5ac15ae7c76053665bbec5ac8a4d Reproduction Job: https://huggingface.co/jobs/visv-Bro/6a6f72c86b79c09949c1f7f1 Job staging Bucket:… See the full description on the dataset page: https://huggingface.co/datasets/visv-Bro/brain-ai-alignment-reproduction.imagen<1K0 likes166 downloads2mo agoHugging Face20ivezakis /llava_med_alignment_500k_chunk_1image10K<n<100K0 likes158 downloads10mo agoHugging Face21ivezakis /llava_med_alignment_500k_chunk_3image10K<n<100K0 likes152 downloads10mo agoHugging Face22china-sae-robotics /camera-calib-and-scene-alignment-dataimage0 likes130 downloads3mo agoHugging Face23ivezakis /llava_med_alignment_500k_chunk_4image10K<n<100K0 likes117 downloads10mo agoHugging Face24Shubhangi29 /llava_med_alignment_500k_chunk_1image100K<n<1M2 likes103 downloads2y agoHugging Face25Rapidata /sora-video-generation-alignment-likert-scoring Rapidata Video Generation Prompt Alignment Dataset If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~6000 human evaluators were asked to evaluate AI-generated videos based on how well the generated video matches the prompt. The specific question… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-alignment-likert-scoring.imagevideo-classificationn<1K18 likes103 downloads2y agoHugging Face26vaibhavalakshmiravideshik /mesh-snomed-entity-alignment-15k MeSH-SNOMED Entity Alignment 15K MeSH-SNOMED Entity Alignment 15K is a biomedical heterogeneous knowledge graph alignment benchmark for cross-ontology matching between MeSH and SNOMED CT. It is designed to evaluate entity alignment systems under realistic large-graph conditions, where gold-aligned concepts are embedded in much larger biomedical graphs containing many structurally relevant but non-aligned background entities. This release is intended for the accompanying EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavalakshmiravideshik/mesh-snomed-entity-alignment-15k.image10K<n<100K2 likes94 downloads5mo agoHugging Face27ivezakis /llava_med_alignment_500k_chunk_5image10K<n<100K0 likes89 downloads10mo agoHugging Face28Project-AgML /Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset Agri-LLaVA Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below). This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.imageimage-text-to-text100K<n<1M0 likes79 downloads3mo agoHugging Face29Rapidata /117k_human_alignment_flux1.0_V_flux1.1Blueberry Rapidata Image Generation Alignment Dataset This Dataset is a 1/3 of a 340k human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/117k_human_preferences_flux1.0_V_flux1.1Blueberry Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/117k_human_coherence_flux1.0_V_flux1.1Blueberry It was collected in ~2 Days using the Rapidata… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/117k_human_alignment_flux1.0_V_flux1.1Blueberry.image1K<n<10K11 likes60 downloads2y agoHugging Face30AlignmentLab-AI /gpt4v-raw-chunksimage100K<n<1M0 likes55 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.