Team Ai
Datasetpublic

houlab/cossmos-annotations-db

ARSMA-web motif annotations (cossmos-annotations-db) Per-motif annotation: base pairs, stacking, sugar puckers, glycosidic conformations and the deposition metadata of the parent entry. It backs the motif browser of ARSMA-web. This is not the occurrence count. One row of instances.parquet is one annotated motif site, and the table has 283,164 of them against 285,165 clips in houlab/motif-db. The two differ because the annotation rows of the 25 CoSSMos classes are CoSSMos's own… See the full description on the dataset page: https://huggingface.co/datasets/houlab/cossmos-annotations-db.

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes257downloads
Dataset Card

ARSMA-web motif annotations (cossmos-annotations-db)

Per-motif annotation: base pairs, stacking, sugar puckers, glycosidic conformations and the deposition metadata of the parent entry. It backs the motif browser of ARSMA-web.

This is not the occurrence count. One row of instances.parquet is one annotated motif site, and the table has 283,164 of them against 285,165 clips in houlab/motif-db. The two differ because the annotation rows of the 25 CoSSMos classes are CoSSMos's own export, which does not match the clip set exactly (2,001 more clips than export rows in those classes, net). To count motif occurrences, count rows in motif-db; use this table for what the motifs contain.

Files

fileone row isnotes
instances.parquetone annotated motif sitesource is cossmos (269,735 rows, the CoSSMos export) or arsma-dssr (13,429 rows, the 19 classes ARSMA-web adds, annotated with DSSR from its own clips)
residues.parquetone residue of a sitesugar pucker, glycosidic angle class; joins on instance_id
basepairs.parquetone base pair or stack of a siteLeontis–Westhof and Saenger classes; joins on instance_id
entries.parquetone PDB entrytitle, method, resolution, organisms, Rfam and RNAcentral cross-references
groups.parquetone (class, sequence) grouphow many sites share the sequence and how many geometric clusters they form
xrefs*.parquetone entry or chainexternal cross-references used to fill entries
facets.json, quality.json—precomputed filter values; build and link statistics

Joining to the other datasets

clip_stem is the motif-db identifier of the clip at the same position (clip_link_status says how it was matched, or missing). map_id/emdb_id point to the density segment in motif-cryomap-db, cluster_id to motif-cluster-db.

Citation

RNA CoSSMos (Vanegas et al., Nucleic Acids Res. 2012; Richardson, Kirkpatrick & Znosko, Database 2020); annotations parsed and normalized by ARSMA-web. DSSR: Lu, Bussemaker & Olson, Nucleic Acids Res. 2015.