Team Ai
Datasetpublic

FronyAI/ko-multisource-retrieval-dataset

구성 여러 공개 출처를 하나의 스키마로 통합한 한국어 query–passage 페어 데이터셋입니다. stage1, stage2 두 subset. 각각 train 100,000 / valid 10,000행 (총 220,000행, parquet, 151MB) 필드 설명 dataset 원 출처 식별자 (18종) query 질의문 (2~2,020자) passage 대응 문서 (10~2,050자) 출처 AI 허브 — 기계독해 지식검색 일반상식 도서자료 행정문서 뉴스기사 금융법률 숫자연산 표정보 추상요약 이벤트 (AI허브_ 접두사) KLUE — klue-mrc klue-nli klue-sts 기타 — kakao-nli 공공데이터포털-deepqa LGNLP 참고 출처별로 페어 성격이 다르므로 용도에 맞게 필터링을 권장합니다. klue-nli klue-sts… See the full description on the dataset page: https://huggingface.co/datasets/FronyAI/ko-multisource-retrieval-dataset.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes20downloads
settings

This repository belongs to FronyAI on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameko-multisource-retrieval-dataset
visibilitypublic
licencenot set
gatedno
ownerFronyAI
Account settings
FronyAI/ko-multisource-retrieval-dataset · Team Ai