Team Ai
Datasetpublic

proteinglm/contact_prediction_binary

Dataset Card for Contact Prediction Dataset Dataset Summary Contact map prediction aims to determine whether two residues, $i$ and $j$, are in contact or not, based on their distance with a certain threshold ($<$8 Angstrom). This task is an important part of the early Alphafold version for structural prediction. Dataset Structure Data Instances For each instance, there is a string of the protein sequences, a sequence for the contact… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/contact_prediction_binary.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes139downloads
Dataset Card

Dataset Card for Contact Prediction Dataset

Dataset Summary

Contact map prediction aims to determine whether two residues, $i$ and $j$, are in contact or not, based on their distance with a certain threshold ($<$8 Angstrom). This task is an important part of the early Alphafold version for structural prediction.

Dataset Structure

Data Instances

For each instance, there is a string of the protein sequences, a sequence for the contact labels. Each of the sub-labels "[2, 3]" indicates the 3rd residue are in contact with the 4th residue (start from index 0). See the Contact map prediction dataset viewer to explore more examples.

{'seq':'QNLLKNLAASLGRKPFVADKQGVYRLTIDKHLVMLAPHGSELVLRTPIDAPMLREGNNVNVTLLRSLMQQALAWAKRYPQTLVLDDCGQLVLEARLRLQELDTHGLQEVINKQLALLEHLIPQLTP'
'label': [ [ 0, 0 ], [ 0, 1 ], [ 1, 1 ], [ 1, 2 ], [ 1, 3 ], [ 1, 101 ], [ 2, 2 ], [ 2, 3 ], [ 2, 4 ], [ 3, 3 ], [ 3, 4 ], [ 3, 5 ], [ 3, 99 ], [ 3, 100 ], [ 3, 101 ], [ 4, 4 ], [ 4, 5 ], [ 4, 53 ], ...]}

The average for the seq and the label are provided below:

FeatureMean Count
seq249
label1,500

Data Fields

  • —seq: a string containing the protein sequence
  • —label: a string containing the contact label of each residue pair.

Data Splits

The contact map prediction dataset has 3 splits: train, validation, and test. Below are the statistics of the dataset.

Dataset SplitNumber of Instances in Split
Train12,041
Validation1,505
Test1,505

Source Data

Initial Data Collection and Normalization

The trRosetta dataset is employed as the initilized dataset.

Licensing Information

The dataset is released under the Apache-2.0 License.

Citation

If you find our work useful, please consider citing the following paper:

@misc{chen2024xtrimopglm,
  title={xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein},
  author={Chen, Bo and Cheng, Xingyi and Li, Pan and Geng, Yangli-ao and Gong, Jing and Li, Shen and Bei, Zhilei and Tan, Xu and Wang, Boyan and Zeng, Xin and others},
  year={2024},
  eprint={2401.06199},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  note={arXiv preprint arXiv:2401.06199}
}