Team Ai
Datasetpublic

yuyijiong/Multi-doc-QA-CommonCrawl

Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first. English multi document Q&A data created using RedPajamaCommonCrawl data as reference text In the Raw dataset, each sample contains one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document. It can be used to train models to extract the target information from a large number of documents. After filtering, integrating, and… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Multi-doc-QA-CommonCrawl.

sourceHugging Facecc-by-nc-4.0updated 3y agoView on Hugging Face
9likes55downloads
Dataset Card
  • —Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first.
  • —English multi document Q&A data created using RedPajamaCommonCrawl data as reference text
  • —In the Raw dataset, each sample contains <font color=red> one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document</font>. It can be used to train models to extract the target information from a large number of documents.
  • —After filtering, integrating, and transforming the raw data into chatml format instruction fine-tuning data, each sample contains approximately 30 reference documents and 5 corresponding QA pairs.

<br/>

  • —2023.12.4更新:改进答案的格式,强制所有答案在回答时必须先给出原文。
  • —以RedPajamaCommonCrawl数据为参考文本,制作的英文多文档问答数据
  • —原始数据中,每个样本包含 <font color=red> 一个参考文档、199个无关文档、一个基于参考文档的问答对</font>。可以训练模型从大量文档中抽取关键信息的能力。
  • —原始数据经过筛选、整合转化为chatml形式的指令微调数据后,每条数据大约包含30个参考文档,以及5个对应的问答对。
  • —dataset size: 11k