ms2
Datasets
All datasets matching “ms2”ms2prospect-ptms-ms2
PROSPECT PTMs - Fragment Ion Intensity Prediction (MS2)
A mass-spectrometry dataset for applied machine learning in proteomics, annotated, processed and split for the task of fragment ion intensity prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo [3][4][5][6].
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-ms2.massive-v2-ms2-t095-l080-sharded-10gb
MassIVE v2 exact-MS2 training shards
This dataset is a training-oriented repack of
novogaia/massive-v2 at
revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in
_t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2.
Rows: 1,584,408,553
Eligible training rows: 1,516,329,213
Train shards: 65
Validation shards: 3
The training_eligible column records the canonical precursor, retention-time,
and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.black-cube-v1
black-cube-v1
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
massive-v1-ms2-100m-stratified-x16
MassIVE v1 MS2 100M Stratified HDF5 Shards
This dataset contains a 100,000,000-row chunk-aligned stratified sample from the MassIVE v1 MS2 HDF5 handoff file.
Source object: gs://metal-repeater-411410-spectra-checkpoints/test_data_massive/v1_handoff/MassIVE_v1_ms2.hdf5
Source rows: 1,620,804,459
Sampled rows: 100,000,000
Shards: 16 HDF5 files, 6,250,000 rows per shard
Compression: HDF5 gzip level 1
Sampling unit: source HDF5 row chunk, 256 rows per source chunk
Stratification… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v1-ms2-100m-stratified-x16.ms2_dense_maxThis is a copy of the MS^2 dataset, except the input source documents of its validation split have been replaced by a dense retriever. The retrieval pipeline used:
query: The background field of each example
corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract.
retriever: facebook/contriever-msmarco via PyTerrier with default settings
top-k strategy: "max", i.e. the number of documents retrieved, k, is set as… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ms2_dense_max.
