Anoy123423123/MSA_PretrainData
MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.
MSA Pretrain Data
Retrieval-style pretraining corpora. Each subset is split into two parts:
reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's references/ table, taken as the concatenation of its parquet shards in filename order.
