proteinglm/enzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset Dataset Summary This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/enzyme_catalytic_efficiency.
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $kcat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $kcat$ predictions applicable to any organism or enzyme. Such capabilities are crucial for advancing our understanding of cellular metabolism and physiology.
Dataset Structure
Data Instances
For each instance, there is a string representing the protein sequence and a float value indicating the $kcat$ score of the protein sequence. See the [ enzyme catalytic efficiency dataset viewer](https://huggingface.co/datasets/Bo1015/enzymecatalytic_efficiency/viewer) to explore more examples.
{'seq':'MEHVIDNFDNIDKCLKCGKPIKVVKLKYIKKKIENIPNSHLINFKYCSKCKRENVIENL'
'label':3.6}The average for the seq and the label are provided below:
Data Fields
seq: a string containing the protein sequencelabel: a float value indicating the $k_cat$ score of the protein sequence.
Data Splits
The Enzyme Catalytic Efficiency dataset has 3 splits: train, valid and test. Below are the statistics of the dataset.
Source Data
Initial Data Collection and Normalization
The data, sourced from a variety of repositories including BRENDA, SABIO-RK, KEGG, UniProt, and MetaCyc, are curated by Li et al.
Licensing Information
The dataset is released under the Apache-2.0 License.
Citation
If you find our work useful, please consider citing the following paper:
@misc{chen2024xtrimopglm,
title={xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein},
author={Chen, Bo and Cheng, Xingyi and Li, Pan and Geng, Yangli-ao and Gong, Jing and Li, Shen and Bei, Zhilei and Tan, Xu and Wang, Boyan and Zeng, Xin and others},
year={2024},
eprint={2401.06199},
archivePrefix={arXiv},
primaryClass={cs.CL},
note={arXiv preprint arXiv:2401.06199}
}