ternarytx/flock
Flock: A negative-enriched protein-protein interaction dataset - Dataset Card Code: https://github.com/TernaryTx/flock Overview Files file rows contents flock_pairs 404,577 the full dataset: one row per unordered pair, with label, source and leakage flags and examples leakage_free_pairs 16,210 a benchmark for cofolding models: 335 targets, leakage-filtered leaked_pairs 14,402 298 targets whose pairs are leaked, as a contrast set… See the full description on the dataset page: https://huggingface.co/datasets/ternarytx/flock.
Flock: A negative-enriched protein-protein interaction dataset - Dataset Card
Code: https://github.com/TernaryTx/flock
Overview
Files
`flock_pairs` is ordered so protein_a < protein_b alphabetically, and no pair appears twice. 26,934 positive and 377,643 negative pairs. 8,310 pairs from PINDER alone, 4,980 from PPI3D alone, and 13,644 pairs from both (indicated in the positive_source column). negative_source is pdb for 368,937 and literature for 8,706. Each label's source column is empty on the other label. Gene names are present for 383,801 of the 404,577 pairs. The four leakage columns are described under Leakage.
`leakage_free_pairs` and `leaked_pairs` carry target, partner and type only. They are directed; the orientation of pairs is meaningful since we calculate per-target metrics. Everything else about a pair lives in flock_pairs and can be joined on the unordered pair.
`benchmark_targets` gives min_num_chains and max_num_chains over the target's source complexes, min_buried_sasa and max_buried_sasa in Ų, and one representative example_pdb so a reader can look at a real structure. Gene names are present for 546 of the 610 targets.
`literature_negatives_detail` holds 10,480 rows over 10,082 pairs. A pair supported by two papers gets a row for each. Alongside the accessions it contains the source paper (pmid, pmcid, text_source), the sentence relied on and where it appeared (evidence_sentence, evidence_location), the model's call and its stated reasoning, and the curation verdict.
usable_as_negativeandpublished_as_negativeare not the same. Curation accepting a pair is not the last word — a final filter drops some of them. 8,896 rows, covering all 8,706 published literature negatives, arepublished_as_negative; the other 1,584 rows describe pairs that did not make it. Every published literature negative has at least one evidence row.
`sequences` gives accession, sequence and length for the two benchmark sets.
`esmfold2_metrics` contains ESMFold2-Fast [1] scores for both benchmark sets in one file, with the benchmark set labelled in the set column: 10,573 pairs over 318 targets for leakage_free, and 13,449 pairs over 298 targets for leaked. It contains ptm, iptm, avg_plddt, pair_chains_iptm, the interface-restricted iptm_d0chn_A_P and three ipSAE [2] variants, alongside n_residues and predict_seconds.
`esmc_metrics` contains ESMC [1] scores for the same pairs as esmfold2_metrics, with the ESMC features normalized over two different feature corpora (see below). 48,044 rows in total. Its score column is npmi_score, alongside set, feature_corpus, method, target, partner, pair_key, n_residues, type, status and imputed_zero. status is OK on every row.
ESMC scores a pair from the normalised pointwise mutual information (NPMI) between the sparse autoencoder features of the two proteins, counted over a corpus of PDB structures [1]. We provide ESMC scores normalized against two different feature corpora, both drawn from PPI3D [6] and labelled in feature_corpus:
full— the corpus up to 2026-07-29, 631,509 feature pairs.pre_2023-06-01_strict— the corpus restricted to structures released before the benchmark cutoff, 478,830 feature pairs. This allows fair comparisons to ESMFold2 on the Leakage-Free set. The date cut is taken per structure cluster rather than per structure: PPI3D clusters interfaces at 40% sequence identity and 50% interface contact similarity, and any cluster holding a member released on or after the cutoff is dropped whole rather than falling back to its best surviving pre-cutoff member, so no interface family left in the corpus grew after the cutoff.
Each corpus covers the exact same 24,022 pairs. 426 pairs per corpus (188 leakage_free and 238 leaked) are scored 0 and marked imputed_zero: ESMC found no residue in one of the two proteins whose features passed the activation threshold.
`zhang_pairs`, `zhang_sequences` and `zhang_esmc_metrics` are an external test set and our result on it, they share no pairs, targets or leakage flags with any file above.
zhang_pairs is the human PPI control set of Zhang et al. [15], parsed from their Data S1. Their positives are 3,000 pairs sampled from the 3,988 confident PPIs that STRING, BioGRID and UniProt all agree on; their negatives are 30,000 random human protein pairs with no evidence of interaction. We then filtered that set for homology against the PPI3D [6] structures that fed into our NPMI feature table, using MMseqs2 clustering [16] at 30% sequence identity and 80% coverage. 6,409 of its 17,266 proteins matched the corpus at that threshold, taking it from 32,747 pairs (3,000 positive and 29,747 negative) to the 12,635 published here: 738 positive and 11,897 negative over 9,820 proteins. It carries protein_a, protein_b and type, with the accessions sorted as in flock_pairs.
zhang_sequences gives accession, sequence and length for all 9,820 proteins, the exact sequences that were scored.
zhang_esmc_metrics holds ESMC's NPMI score over that set against the full feature corpus, one row per pair. Alongside npmi_score it carries the per-side active feature counts n_active_a, n_active_b, n_selected and n_positive, plus status and imputed_zero. Pooled AUROC is 0.9015 imputing the 67 no-feature pairs to zero, or 0.9010 dropping them.
Reproduction note: We endeavoured to reproduce the published ESMC method as best as possible; however some details and parameters were not described in their paper. In particular the exact date cutoff behind their feature corpus is not provided; our PPI3D cutoff of 2026-07-29 may be slightly later than theirs and as such our NPMI set may be slightly bigger. In our paper we report comparable performance on this set.
Loading
To train models on Flock, load the data using HuggingFaceDatasets:
from datasets import load_dataset
pairs = load_dataset("ternarytx/flock", "flock_pairs", split="train")Other configurations are leakage_free_pairs, leaked_pairs, benchmark_targets, literature_negatives_detail, sequences, esmfold2_metrics, esmc_metrics, zhang_pairs, zhang_sequences and zhang_esmc_metrics.
print(pairs)
# Dataset({
# features: ['protein_a', 'protein_b', 'gene_a', 'gene_b', 'label',
# 'positive_source', 'negative_source', 'in_training_set',
# 'homolog_in_training_set', 'in_training_set_entry',
# 'homolog_in_training_set_entry'],
# num_rows: 404577
# })
print(pairs[0])
# {'protein_a': 'A0A009I821', 'protein_b': 'A0A1V3DIZ9',
# 'gene_a': 'rplV', 'gene_b': 'rpsP',
# 'label': 'Negative', 'positive_source': None, 'negative_source': 'pdb',
# 'in_training_set': True, 'homolog_in_training_set': True,
# 'in_training_set_entry': '7M4W', 'homolog_in_training_set_entry': '7M4W'}
from collections import Counter
print(Counter(pairs["label"]))
# Counter({'Negative': 377643, 'Positive': 26934})
print(Counter(pairs["negative_source"]))
# Counter({'pdb': 368937, None: 26934, 'literature': 8706})
# "None" pairs are positivesThe CSVs can also be read directly for analysis and plots:
import pandas as pd
from huggingface_hub import hf_hub_download
def flock_csv(name: str) -> pd.DataFrame:
return pd.read_csv(
hf_hub_download("ternarytx/flock", name, repo_type="dataset"),
)
flock = flock_csv("flock_pairs_v1_2026-08-24.csv")
benchmark = flock_csv("leakage_free_pairs_v1_2026-08-24.csv")
sequences = flock_csv("sequences_v1_2026-08-24.csv")leakage_free_pairs and leaked_pairs are keyed on target and partner, while flock_pairs is keyed on unordered pairs. To join them, first sort on the two accessions:
# Sort alphabetically to match Flock pairs
benchmark["protein_a"] = benchmark[["target", "partner"]].min(axis=1)
benchmark["protein_b"] = benchmark[["target", "partner"]].max(axis=1)
annotated = benchmark.merge(flock, on=["protein_a", "protein_b"], how="left")To attach sequences to a benchmark pair, map each side in turn:
lookup = sequences.set_index("accession")["sequence"]
benchmark["target_sequence"] = benchmark["target"].map(lookup)
benchmark["partner_sequence"] = benchmark["partner"].map(lookup)Leakage
A pair in Flock is labelled as leaked if it (or any homologous pair) could have appeared in the training set of a cofolding model. The Boltz-2 [3] cutoff date of 2023-06-01 is used, which also covers any model matching the AlphaFold 3 [4] training cutoff date (OpenFold3, ESMFold2, Protenix, Chai-1 etc.).
Leakage Flags
in_training_setdenotes pairs where a pre-cutoff PDB entry contains both exact proteins: the complex itself for a positive, or co-presence in one entry for a negative.- Positives 21,244 True / 5,690 False. Negatives 223,227 True / 154,416 False.
homolog_in_training_setdenotes pairs where a similar pair exists in a pre-cutoff PDB entry. Interface-cluster match to a pre-cutoff complex for positives; sequence identity at 40% or above to a pre-cutoff pair for negatives.- Positives 22,204 True / 4,730 False. Negatives 283,232 True / 94,411 False.
- Clean on both flags: 3,275 positives and 86,827 negatives.
Only pairs with both flags False are kept. Then, we filter to a set of targets such that each target has at least 2 clean positive and 20 clean negative partners.
This gives leakage_free_pairs, 16,210 rows over 335 targets.
Pre-Cutoff Entry Examples
flock_pairs contains columns in_training_set_entry and homolog_in_training_set_entry. Each names a PDB entry that supports the flag.
They are populated for negatives only. A negative's flags are decided per pair, against a specific PDB entry: the one putting both proteins in a single pre-cutoff structure, or the one holding the homologous pair. Every negative pair has at least one example leaked entry.
A positive's flags are decided against a set; the release dates of all its source complexes for in_training_set, and membership of an interface cluster for homolog_in_training_set. As such there is no single entry to point at and these two example entry columns are unpopulated.
Leaked Pairs
The leaked_pairs set is a contrast dataset built from pairs that failed the leakage filter, matched on length distribution and with the same filters (at least 2 positives and 20 negatives) as the benchmark experiment ran on the leakage free set. This is so that scores can be compared against pairs a model may have memorised (see paper for details). The scores behind that comparison are in esmfold2_metrics and esmc_metrics, each holding both sets.
*Note: a protein can appear in both leakage-free and leaked sets. Leakage is a property of pairs, not individual proteins. For example, A8JFK6 has 64 pairs in Flock, 33 clean and 31 leaked, so it is a clean target with 33 partners and a leaked target with 31 different ones. No pair appears in both leakage-free and leaked sets*.
Sources
The access dates bound what could have leaked, so they are given rather than implied. The PPI3D pull is eight months later than the PINDER snapshot, and supplies the structures released in between.
Licence
Flock is released under CC BY 4.0. Use it for anything, including commercially, provided you credit this dataset and its sources by citing the paper and the references below.
The sources carry their own terms, all of which permit redistribution here:
The pairs in zhang_pairs and zhang_esmc_metrics are a filtered subset of the control sets in Data S1 of Zhang et al. [15], redistributed here with attribution so that the scores published beside them can be checked.
Two additional sources were used to generate this dataset but do not appear in it. IntAct [8] is used only as a filter, to remove pairs with curated evidence of interaction from the negatives, so no IntAct records are redistributed here. Negatome 2.0 [14] inspired this paper, and while this work builds on the methods used there, the data presented here is generated from different inputs, sources, and code.
literature_negatives_detail quotes short passages from the articles it cites by PMID [9]. Rights in that quoted text remain with the original publishers.
References
- Candido, S. et al. Language Modeling Materializes a World Model of Protein Biology. 2026.06.03.729735 Preprint at https://doi.org/10.64898/2026.06.03.729735 (2026).
- Dunbrack, R. L. Rēs ipSAE loquuntur: What's wrong with AlphaFold's ipTM score and how to fix it. 2025.02.10.637595 Preprint at https://doi.org/10.1101/2025.02.10.637595 (2025).
- Passaro, S. et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. bioRxiv 2025.06.14.659707 (2025) doi:10.1101/2025.06.14.659707.
- Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).
- Kovtun, D. et al. PINDER: The protein interaction dataset and evaluation resource. 2024.07.17.603980 Preprint at https://doi.org/10.1101/2024.07.17.603980 (2024).
- Dapkūnas, J., Timinskas, A., Olechnovič, K., Tomkuvienė, M. & Venclovas, Č. PPI3D: a web server for searching, analyzing and modeling protein–protein, protein–peptide and protein–nucleic acid interactions. Nucleic Acids Res 52, W264–W271 (2024).
- Berman, H. M. et al. The Protein Data Bank. Nucleic Acids Res 28, 235–242 (2000).
- Orchard, S. et al. The MIntAct project—IntAct as a common curation platform for 11 molecular interaction databases. Nucleic Acids Res 42, D358–D363 (2014).
- Rosonovski, S. et al. Europe PMC in 2023. Nucleic Acids Res 52, D1668–D1676 (2024).
- Kafkas, Ş. et al. Section level search functionality in Europe PMC. J Biomed Semant 6, 7 (2015).
- The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res 53, D609–D617 (2025).
- Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
- Varadi, M. et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res 52, D368–D375 (2024).
- Blohm, P. et al. Negatome 2.0: a database of non-interacting proteins derived by literature mining, manual annotation and protein structure analysis. Nucleic Acids Res 42, D396-400 (2014).
- Zhang, J. et al. Predicting protein-protein interactions in the human proteome. Science 390, eadt1630 (2025).
- Steinegger, M. & Söding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35, 1026–1028 (2017).
