macwiatrak/bacbench-ppi-stringdb-dna-small
Dataset for protein-protein interaction prediction across bacteria (DNA) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/). Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores. The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.
Dataset for protein-protein interaction prediction across bacteria (DNA)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/). Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores. The PPI scores have been extracted using the combined score from STRING DB.
The interaction between two proteins is represented by a triple: [prot1_index, prot2_index, score]. Where to get a probability score, you must divide the score by 1000 (i.e. if the score is 721 then to get a true score do 721/1000=0.721). The index of a protein refers to the index of the protein in the protein_sequences column of the row. See example below in Usage
Usage
We recommend loading the dataset in a streaming mode to prevent memory errors.
from datasets import load_dataset
ds = load_dataset("macwiatrak/bacbench-ppi-stringdb-dna-small", split="validation", streaming=True)
item = next(iter(ds))
# select a contig_idx and gene_idx
contig_idx = 0
gene_idx = 0
# fetch protein sequences from a genome (list of strings) for the contig_idx
dna_seq = item["dna_sequence"][contig_idx]
# get gene sequence, indices are 1-based inclusive so we account for it
start_idx = item['start'][contig_idx][gene_idx] - 1
end_idx = item['end'][contig_idx][gene_idx]
strand_idx = item['strand'][contig_idx][gene_idx] # we can also get strand which can be 1 (positive) or -1 (negative)
gene_seq = dna_seq[start_idx:end_idx]
# fetch PPI triples labels (i.e. [prot1_index, prot2_index, score])
ppi_triples = item["labels"][contig_idx]
# get protein seqs and label for one pair of proteins
prot1 = prot_seqs[ppi_triples[0][0]]
prot2 = prot_seqs[ppi_triples[0][1]]
score = ppi_triples[0][2] / 1000
# we recommend binarizing the labels based on the threshold of 0.6
binary_ppi_triples = [
(prot1_index, prot2_index, int((score / 1000) >= 0.6)) for prot1_index, prot2_index, score in ppi_triples
]Split
We provide a phylogeny-aware train, validation and test split by genus with proportions of 60 / 10 / 20 (%) respectively as part of the dataset. This means that the the genera in train, validation and test do not overlap.
See github repository for details on how to embed the dataset with DNA and protein language models as well as code to predict protein-protein interactions.
Other relevant resources:
- Equivalent dataset with protein sequences rather than DNA - https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small
- Full dataset of bacterial organisms with associated PPI from STRING DB (10,533 genomes) and protein sequences (not DNA) - https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences
