Team Ai
Datasetpublic

proteinglm/enzyme_catalytic_efficiency

Dataset Card for Enzyme Catalytic Efficiency Dataset Dataset Summary This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/enzyme_catalytic_efficiency.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes32downloads
Dataset Card

Dataset Card for Enzyme Catalytic Efficiency Dataset

Dataset Summary

This task is focused on predicting $kcat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $kcat$ predictions applicable to any organism or enzyme. Such capabilities are crucial for advancing our understanding of cellular metabolism and physiology.

Dataset Structure

Data Instances

For each instance, there is a string representing the protein sequence and a float value indicating the $kcat$ score of the protein sequence. See the [ enzyme catalytic efficiency dataset viewer](https://huggingface.co/datasets/Bo1015/enzymecatalytic_efficiency/viewer) to explore more examples.

{'seq':'MEHVIDNFDNIDKCLKCGKPIKVVKLKYIKKKIENIPNSHLINFKYCSKCKRENVIENL'
'label':3.6}

The average for the seq and the label are provided below:

FeatureMean Count
seq418
label1.87

Data Fields

  • —seq: a string containing the protein sequence
  • —label: a float value indicating the $k_cat$ score of the protein sequence.

Data Splits

The Enzyme Catalytic Efficiency dataset has 3 splits: train, valid and test. Below are the statistics of the dataset.

Dataset SplitNumber of Instances in Split
Train13,470
Valid1,684
Test1,684

Source Data

Initial Data Collection and Normalization

The data, sourced from a variety of repositories including BRENDA, SABIO-RK, KEGG, UniProt, and MetaCyc, are curated by Li et al.

Licensing Information

The dataset is released under the Apache-2.0 License.

Citation

If you find our work useful, please consider citing the following paper:

@misc{chen2024xtrimopglm,
  title={xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein},
  author={Chen, Bo and Cheng, Xingyi and Li, Pan and Geng, Yangli-ao and Gong, Jing and Li, Shen and Bei, Zhilei and Tan, Xu and Wang, Boyan and Zeng, Xin and others},
  year={2024},
  eprint={2401.06199},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  note={arXiv preprint arXiv:2401.06199}
}