yuyijiong/Multi-doc-QA-CommonCrawl
Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first. English multi document Q&A data created using RedPajamaCommonCrawl data as reference text In the Raw dataset, each sample contains one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document. It can be used to train models to extract the target information from a large number of documents. After filtering, integrating, and… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Multi-doc-QA-CommonCrawl.
955
- Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first.
- English multi document Q&A data created using RedPajamaCommonCrawl data as reference text
- In the Raw dataset, each sample contains <font color=red> one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document</font>. It can be used to train models to extract the target information from a large number of documents.
- After filtering, integrating, and transforming the raw data into chatml format instruction fine-tuning data, each sample contains approximately 30 reference documents and 5 corresponding QA pairs.
<br/>
- 2023.12.4更新:改进答案的格式,强制所有答案在回答时必须先给出原文。
- 以RedPajamaCommonCrawl数据为参考文本,制作的英文多文档问答数据
- 原始数据中,每个样本包含 <font color=red> 一个参考文档、199个无关文档、一个基于参考文档的问答对</font>。可以训练模型从大量文档中抽取关键信息的能力。
- 原始数据经过筛选、整合转化为chatml形式的指令微调数据后,每条数据大约包含30个参考文档,以及5个对应的问答对。
- dataset size: 11k
