Team Ai
Datasetpublic

Anoy123423123/MSA_PretrainData

MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes4kdownloads
Dataset Card

MSA Pretrain Data

Retrieval-style pretraining corpora. Each subset is split into two parts:

filecolumnsmeaning
<subset>/queries/*.parquetquestion, answer, reference_ids: list<int64>, labels: list<int64>query, plus row indices into the subset's reference table
<subset>/references/*.parquetvalue: stringthe reference/memory passage text

reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's references/ table, taken as the concatenation of its parquet shards in filename order.