datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sea-commoncrawlsea-commoncrawl-high-qualitycitations
Common Crawl Citations Overview
This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar.
Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations.
CommonCrawl-RAG-QA-Calm3-22b-chat
自動生成テキスト
データソースから、OpenCalm3-22bを使ってクリーニング・再生成したテキストです。
Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
CommonLID
CommonLID: Language Identification for the Real Web
CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data.
The number of lines available for each language is provided in Appendix A of the paper.
➡️ CommonLID Leaderboard on Hugging Face 🤗
➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.nigeriaai-commoncrawlicelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
commoncrawl-jobs-demo
Common Crawl on Jobs — datatrove JobsPipelineExecutor demo
228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster.
This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.Multi-doc-QA-CommonCrawl
Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first.
English multi document Q&A data created using RedPajamaCommonCrawl data as reference text
In the Raw dataset, each sample contains one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document. It can be used to train models to extract the target information from a large number of documents.
After filtering, integrating, and… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Multi-doc-QA-CommonCrawl.GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl
Description
This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset.
This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl:
id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.
