Team Ai
Datasetpublic

eczech/marinfold-exp11-protein-docs-seq

marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes191downloads
Dataset Card

marinfold-exp11-pdocs-seq

Sequence-only derivative of `eczech/marinfold-exp11-protein-docs`.

For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so it remains compatible with the same tokenizer.

All other columns, the subset (high/medium/low) partitioning, and the train/validation/test split assignment are inherited unchanged from the source dataset.