datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.korean_data_scraper_wikipedia
Korean Data Scraper — Wikipedia Dump Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 wikipedia_dump 소스가 생성한 코퍼스입니다. 한국어 위키백과(ko.wikipedia.org) 덤프를 파싱하여 문서 본문 텍스트를 추출한 결과입니다.
스키마
파일당 1개 레코드(JSONL)이며, korean_data_scraper_kakaotalk와 동일한 공통 스키마를 따릅니다.
{"id": "kowiki:70773", "source": "wikipedia_dump", "text": "..."}
id는 kowiki:<문서 ID> 형태이며, text는 해당 위키백과 문서의 본문 텍스트입니다.
데이터 규모 및 한계
전체 32개 샤드(shard-00000~00031)로 구성되며, 총 용량은 약 2.24GB입니다. 이 중… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_wikipedia.web-scraper-dataset
Web Scraper Dataset
Whole-site crawl of 650+ domains across ~70 categories (news, wikis, fanfiction,
education, NSFW, social, gaming, code forges, papers, shadow libraries and more).
Common Crawl first (with retry/backoff), live crawlers (wget2/Playwright/Selenium/
requests) as fallback, uncapped per-domain page counts.
Files
MASTER_text.parquet — crawled page text
(id, modality, language, source_domain, category, title, url, text, text_length, crawl_date)… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/web-scraper-dataset.discord-messages
Discord Messages Dataset
Description
This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains.
The data is formatted as plain text with one message per line, making it ideal for:
Language model pre-training
Fine-tuning chatbots
Sentiment analysis
Toxicity detection
Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.korean_data_scraper_kakaotalk
Korean Data Scraper — KakaoTalk Export Corpus
© 주식회사 MKD(MKD Inc.) — 모든 권리 보유 (All rights reserved).
본 데이터셋은 주식회사 MKD(MKD Inc.)가 자체적으로 생성·소유한 내부 데이터이며,
외부 공개 라이선스가 부여되지 않은 사내 전용(proprietary) 자료입니다.
Korean_data_scraper 프로젝트의 kakaotalk_export 소스가 생성한 코퍼스입니다.
사용자 본인이 참여자인 KakaoTalk 그룹 채팅방 2곳의 공식 대화 내보내기(대화 내보내기)
파일을 파싱한 결과이며, 웹 스크래핑이 아니라 사용자가 직접 소유한 원본 데이터를
가공한 것입니다.
비공개(Private) 저장소입니다. 원본 채팅에 사내 인프라 정보와 다수 참여자의
실명이 포함되어 있어, 아래 정제 과정을 거쳤더라도 외부 공개를 의도한 데이터셋이
아닙니다.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_kakaotalk.
