datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
hoptoqa_musique_html_index
hoptoqa_musique_html_index
이 dataset repo는 train_html_hotpotqa_index와 train_html_musique_index를 함께 보관합니다.
포함 항목:
data/train_html_hotpotqa_index/
data/train_html_musique_index/
artifacts/hotpot_distractor_symlink_manifest.json
scripts/restore_musique_hotpot_index_symlinks.py
scripts/bundle_musique_hotpot_index_package.py
복구 정책:
train_html_musique_index에 실제 .npy 파일이 이미 있으면 그대로 유지
비어 있는 페이지나 잘못된 symlink만 manifest 기준으로 train_html_hotpotqa_index를 가리키는 symlink로 복원
예시:
python3… See the full description on the dataset page: https://huggingface.co/datasets/SangMin9806/hoptoqa_musique_html_index.hotpotqa_musique_html_imgmlp-board-yuki-archive-html-textTaken from the yuki archive data. Contains the html text from the /mlp/ board up until a few years ago. It sure would be interesting to have a dataset that's organized by each thread in a spreadsheet separated by row with the posts in the thread separated by column, with the post number in each cell.
MCFTable-HTML
