datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PPIRD
PPIRD: Patent-Product Image Retrieval Dataset
PPIRD is the dataset released with the NeurIPS 2025 paper:
Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval
PPIRD is designed for Patent-Product Image Retrieval (PPIR), where a model retrieves relevant patent images from a large patent gallery given a product image query. This setting is useful for studying patent infringement search, open-set image retrieval, cross-domain visual matching, and… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/PPIRD.MMM-PPICyclinA_RXL_PPI_BLOCKER
Cyclin A RxL PPI Blockers — Ligand–Receptor Complexes
Why this target matters. The cyclin A RxL groove is how cyclin–CDK complexes select their substrates, so blocking it offers a substrate-level selectivity that ATP-competitive CDK inhibitors — all competing for the same conserved pocket — cannot reach.
57 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the RxL substrate-recruitment groove of Cyclin A — a shallow protein–protein… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/CyclinA_RXL_PPI_BLOCKER.ppicsios-app-icons
IOS App Icons
Overview
This dataset contains images and captions of iOS app icons obtained from the iOS Icon Gallery. Each image is paired with a generated caption using a Blip Image Captioning model. The dataset is suitable for image captioning tasks and can be used to train and evaluate models for generating captions for iOS app icons.
Images
The images are stored in the 'images' directory, and each image is uniquely identified with a filename (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ppierzc/ios-app-icons.ecoli_holdout_ppi_large
Clustered PPI datasets (BIOGRID + STRING) with sequence-disjoint splits
This dataset repo contains multiple dataset variants of protein–protein interactions (PPIs),
built by clustering proteins by sequence similarity and then constructing train/valid/test splits that are
intended to be disjoint at the protein level (and thus hard to memorize via near-identical sequences).
Artifacts are stored as compressed pickles (*.pkl.gz). A helper downloader exists in this repo:… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ecoli_holdout_ppi_large.CRASH_Benchmark_CTACF-MS_Homo_sapiens_PPI
CF-MS Elution Profile PPI Dataset
Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.Gutenberg-Fictionatom3d-ppi
PPI: Protein-Protein Interfaces
Overview
This task relates to predicting which pairs of amino acids, spanning two
different proteins, will interact upon binding (when they form a complex).
Amino acids are defined as interacting if any of their heavy atoms are within 6
Angstroms from one another.
Datasets
splits:
DIPS-split: DIPS dataset, split by sequence identity (see add. inf.)
Format
Each entry in the dataset contains the following keys:… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/atom3d-ppi.PhysicalAI-NuRec-PPISP
PPISP Dataset
Dataset Description:
The PPISP dataset accompanies the work "PPISP: Physically-Plausible Compensation and Control of Photometric Variations in Radiance Field Reconstruction". It contains object-centric scene captures of four outdoor scenes, each captured with three different cameras, for multi-view 3D reconstruction and novel view synthesis. The photos were captured with exposure bracketing of +/-2 EV and re-processed with automatic exposure and color… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-NuRec-PPISP.PPLM_PPIplm_interact_human_train_cross_ppiDataset of Human PPI examples with cross-species test examples. Details found here: https://www.nature.com/articles/s41467-025-64512-w
Originally from: https://huggingface.co/datasets/danliu1226/cross_species_benchmarking
Please cite their work.
bacbench-ppi-stringdb-dna-small
Dataset for protein-protein interaction prediction across bacteria (DNA)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/).
Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores.
The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.ppibacbench-ppi-stringdb-protein-sequences-small
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.ppi_SHS148k_bfs_2025Open-PPIyeast-ppibernett_gold_ppi
Leakage-free "gold" standard PPI dataset
From Bernett, et al, found in
Cracking the black box of deep sequence-based protein–protein interaction prediction
paper
code
and
Deep learning models for unbiased sequence-based PPI prediction plateau at an accuracy of 0.65
paper
code
Description
This is a balanced binary protein-protein interaction dataset with positives from HIPPIE and paritioned with KaHIP. There are no sequence overlaps in splits, furthermore, they are… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/bernett_gold_ppi.ppi_dscriptMUSE-benchmark
MUSE: Measuring Uncertainty Source Discrimination
MUSE is a behavioral benchmark designed to evaluate how LLMs distinguish between Epistemic (knowledge gaps) and Aleatoric (stochasticity) uncertainty.
Dataset Summary
This dataset contains 200 items across four dimensions:
E-Type: Pure knowledge gaps.
A-Type: Purely stochastic outcomes.
PA (Pseudo-Aleatoric): Deterministic but complex facts (where the "Trap" occurs).
S (Sycophancy): Adversarial social pressure items.
ppi_SHS148k_dfs_2025artur_studio_ttsTTS slovenian dataset, contains 40 hours of studio recording of a single speaker.
created from:
Verdonik, Darinka; et al., 2023,
ASR database ARTUR 1.0 (audio), Slovenian language resource repository CLARIN.SI, ISSN 2820-4042,
http://hdl.handle.net/11356/1776.
only studio recordings of speaker G0911
recordings without transcriptions were removed
resampled to 22050Hz 16bit wav
metadata.txt contains the transcriptions in format FILENAME_WITHOUT_EXTENSION|SPEAKER_NAME|TRANSCRIPTION
some… See the full description on the dataset page: https://huggingface.co/datasets/ppisljar/artur_studio_tts.ppi_affinityppi_SHS27k_dfs_2025incomplete_ppi_for_speciesPPIs for 160 species out of 300+ in the age dataset.
string_ppi_human_5Msloleks-3-sqlSloleks 3.0 database converted to sqlite format for easier consumption
github repo with scripts and more information: https://github.com/ppisljar/sloleks-3-parser
Citations
Čibej, Jaka; et al., 2022,
Morphological lexicon Sloleks 3.0, Slovenian language resource repository CLARIN.SI, ISSN 2820-4042,
http://hdl.handle.net/11356/1745.
