Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01haifan-gong /PPIRD PPIRD: Patent-Product Image Retrieval Dataset PPIRD is the dataset released with the NeurIPS 2025 paper: Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval PPIRD is designed for Patent-Product Image Retrieval (PPIR), where a model retrieves relevant patent images from a large patent gallery given a product image query. This setting is useful for studying patent infringement search, open-set image retrieval, cross-domain visual matching, and… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/PPIRD.imageimage-to-image10M<n<100M0 likes611 downloads4mo agoHugging Face02yzf1102 /MMM-PPItext1K<n<10K2 likes499 downloads5mo agoHugging Face03Tc-43 /CyclinA_RXL_PPI_BLOCKER Cyclin A RxL PPI Blockers — Ligand–Receptor Complexes Why this target matters. The cyclin A RxL groove is how cyclin–CDK complexes select their substrates, so blocking it offers a substrate-level selectivity that ATP-competitive CDK inhibitors — all competing for the same conserved pocket — cannot reach. 57 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the RxL substrate-recruitment groove of Cyclin A — a shallow protein–protein… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/CyclinA_RXL_PPI_BLOCKER.graph-mln<1K0 likes370 downloads24d agoHugging Face04lordx55 /ppicsimage10K<n<100K0 likes347 downloads6mo agoHugging Face05ppierzc /ios-app-icons IOS App Icons Overview This dataset contains images and captions of iOS app icons obtained from the iOS Icon Gallery. Each image is paired with a generated caption using a Blip Image Captioning model. The dataset is suitable for image captioning tasks and can be used to train and evaluate models for generating captions for iOS app icons. Images The images are stored in the 'images' directory, and each image is uniquely identified with a filename (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ppierzc/ios-app-icons.image1K<n<10K8 likes335 downloads3y agoHugging Face06Synthyra /ecoli_holdout_ppi_large Clustered PPI datasets (BIOGRID + STRING) with sequence-disjoint splits This dataset repo contains multiple dataset variants of protein–protein interactions (PPIs), built by clustering proteins by sequence similarity and then constructing train/valid/test splits that are intended to be disjoint at the protein level (and thus hard to memorize via near-identical sequences). Artifacts are stored as compressed pickles (*.pkl.gz). A helper downloader exists in this repo:… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ecoli_holdout_ppi_large.image0 likes265 downloads8d agoHugging Face07ppino2233 /CRASH_Benchmark_CTAimage100K<n<1M0 likes250 downloads11d agoHugging Face08viridono /CF-MS_Homo_sapiens_PPI CF-MS Elution Profile PPI Dataset Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.text10M<n<100M2 likes229 downloads29d agoHugging Face09macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes227 downloads1y agoHugging Face10ppirli /Gutenberg-Fictiontext10K<n<100K0 likes193 downloads8mo agoHugging Face11vector-institute /atom3d-ppi PPI: Protein-Protein Interfaces Overview This task relates to predicting which pairs of amino acids, spanning two different proteins, will interact upon binding (when they form a complex). Amino acids are defined as interacting if any of their heavy atoms are within 6 Angstroms from one another. Datasets splits: DIPS-split: DIPS dataset, split by sequence identity (see add. inf.) Format Each entry in the dataset contains the following keys:… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/atom3d-ppi.100K<n<1M0 likes131 downloads2y agoHugging Face12nvidia /PhysicalAI-NuRec-PPISP PPISP Dataset Dataset Description: The PPISP dataset accompanies the work "PPISP: Physically-Plausible Compensation and Control of Photometric Variations in Radiance Field Reconstruction". It contains object-centric scene captures of four outdoor scenes, each captured with three different cameras, for multi-view 3D reconstruction and novel view synthesis. The photos were captured with exposure bracketing of +/-2 EV and re-processed with automatic exposure and color… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-NuRec-PPISP.image-to-3d1K<n<10K11 likes116 downloads7mo agoHugging Face13wj5 /PPLM_PPI0 likes98 downloads10mo agoHugging Face14GleghornLab /plm_interact_human_train_cross_ppiDataset of Human PPI examples with cross-species test examples. Details found here: https://www.nature.com/articles/s41467-025-64512-w Originally from: https://huggingface.co/datasets/danliu1226/cross_species_benchmarking Please cite their work. text100K<n<1M0 likes91 downloads1y agoHugging Face15macwiatrak /bacbench-ppi-stringdb-dna-small Dataset for protein-protein interaction prediction across bacteria (DNA) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/). Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores. The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.textn<1K0 likes88 downloads5mo agoHugging Face16Wuming /ppitext10K<n<100K0 likes84 downloads2y agoHugging Face17macwiatrak /bacbench-ppi-stringdb-protein-sequences-small Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.textn<1K0 likes82 downloads5mo agoHugging Face18GleghornLab /ppi_SHS148k_bfs_2025text10K<n<100K0 likes81 downloads1y agoHugging Face19yuyinyang3y /Open-PPItext100K<n<1M1 likes77 downloads1y agoHugging Face20hazemessam /yeast-ppitabular10K<n<100K0 likes56 downloads9mo agoHugging Face21Synthyra /bernett_gold_ppi Leakage-free "gold" standard PPI dataset From Bernett, et al, found in Cracking the black box of deep sequence-based protein–protein interaction prediction paper code and Deep learning models for unbiased sequence-based PPI prediction plateau at an accuracy of 0.65 paper code Description This is a balanced binary protein-protein interaction dataset with positives from HIPPIE and paritioned with KaHIP. There are no sequence overlaps in splits, furthermore, they are… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/bernett_gold_ppi.text100K<n<1M2 likes53 downloads2y agoHugging Face22andyjzhao /ppi_dscripttext100K<n<1M0 likes50 downloads8mo agoHugging Face23Ppilot2 /MUSE-benchmark MUSE: Measuring Uncertainty Source Discrimination MUSE is a behavioral benchmark designed to evaluate how LLMs distinguish between Epistemic (knowledge gaps) and Aleatoric (stochasticity) uncertainty. Dataset Summary This dataset contains 200 items across four dimensions: E-Type: Pure knowledge gaps. A-Type: Purely stochastic outcomes. PA (Pseudo-Aleatoric): Deterministic but complex facts (where the "Trap" occurs). S (Sycophancy): Adversarial social pressure items. question-answering0 likes49 downloads5mo agoHugging Face24GleghornLab /ppi_SHS148k_dfs_2025text10K<n<100K0 likes48 downloads1y agoHugging Face25ppisljar /artur_studio_ttsTTS slovenian dataset, contains 40 hours of studio recording of a single speaker. created from: Verdonik, Darinka; et al., 2023, ASR database ARTUR 1.0 (audio), Slovenian language resource repository CLARIN.SI, ISSN 2820-4042, http://hdl.handle.net/11356/1776. only studio recordings of speaker G0911 recordings without transcriptions were removed resampled to 22050Hz 16bit wav metadata.txt contains the transcriptions in format FILENAME_WITHOUT_EXTENSION|SPEAKER_NAME|TRANSCRIPTION some… See the full description on the dataset page: https://huggingface.co/datasets/ppisljar/artur_studio_tts.text10K<n<100K0 likes47 downloads3y agoHugging Face26Synthyra /ppi_affinitytabular10K<n<100K0 likes41 downloads1y agoHugging Face27GleghornLab /ppi_SHS27k_dfs_2025text1K<n<10K0 likes40 downloads1y agoHugging Face28seq-to-pheno /incomplete_ppi_for_speciesPPIs for 160 species out of 300+ in the age dataset. text1B<n<10B0 likes38 downloads2y agoHugging Face29vladak /string_ppi_human_5Mtabular1M<n<10M1 likes36 downloads1y agoHugging Face30ppisljar /sloleks-3-sqlSloleks 3.0 database converted to sqlite format for easier consumption github repo with scripts and more information: https://github.com/ppisljar/sloleks-3-parser Citations Čibej, Jaka; et al., 2022, Morphological lexicon Sloleks 3.0, Slovenian language resource repository CLARIN.SI, ISSN 2820-4042, http://hdl.handle.net/11356/1745. 0 likes34 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.