ppi
Datasets
All datasets matching “ppi”PPIRD
PPIRD: Patent-Product Image Retrieval Dataset
PPIRD is the dataset released with the NeurIPS 2025 paper:
Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval
PPIRD is designed for Patent-Product Image Retrieval (PPIR), where a model retrieves relevant patent images from a large patent gallery given a product image query. This setting is useful for studying patent infringement search, open-set image retrieval, cross-domain visual matching, and… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/PPIRD.MMM-PPICyclinA_RXL_PPI_BLOCKER
Cyclin A RxL PPI Blockers — Ligand–Receptor Complexes
Why this target matters. The cyclin A RxL groove is how cyclin–CDK complexes select their substrates, so blocking it offers a substrate-level selectivity that ATP-competitive CDK inhibitors — all competing for the same conserved pocket — cannot reach.
57 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the RxL substrate-recruitment groove of Cyclin A — a shallow protein–protein… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/CyclinA_RXL_PPI_BLOCKER.ppicsios-app-icons
IOS App Icons
Overview
This dataset contains images and captions of iOS app icons obtained from the iOS Icon Gallery. Each image is paired with a generated caption using a Blip Image Captioning model. The dataset is suitable for image captioning tasks and can be used to train and evaluate models for generating captions for iOS app icons.
Images
The images are stored in the 'images' directory, and each image is uniquely identified with a filename (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ppierzc/ios-app-icons.ecoli_holdout_ppi_large
Clustered PPI datasets (BIOGRID + STRING) with sequence-disjoint splits
This dataset repo contains multiple dataset variants of protein–protein interactions (PPIs),
built by clustering proteins by sequence similarity and then constructing train/valid/test splits that are
intended to be disjoint at the protein level (and thus hard to memorize via near-identical sequences).
Artifacts are stored as compressed pickles (*.pkl.gz). A helper downloader exists in this repo:… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ecoli_holdout_ppi_large.
