Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sailor2 /sea-commoncrawltext100M<n<1B1 likes4.7k downloads2y agoHugging Face02sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.4k downloads2y agoHugging Face03commoncrawl /citations Common Crawl Citations Overview This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar. Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations. text1K<n<10K5 likes249 downloads6mo agoHugging Face04kanhatakeyama /CommonCrawl-RAG-QA-Calm3-22b-chat 自動生成テキスト データソースから、OpenCalm3-22bを使ってクリーニング・再生成したテキストです。 Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 データ jsonlファイルが数十GB程度あります datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。 text10M<n<100M4 likes198 downloads2y agoHugging Face05commoncrawl /CommonLIDgated CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.texttext-classification100K<n<1M56 likes194 downloads14d agoHugging Face06xerozee /nigeriaai-commoncrawltext1M<n<10M0 likes191 downloads6mo agoHugging Face07mideind /icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3). texttext-generation1M<n<10M1 likes180 downloads4y agoHugging Face08davanstrien /commoncrawl-jobs-demo Common Crawl on Jobs — datatrove JobsPipelineExecutor demo 228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster. This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.tabular100K<n<1M11 likes75 downloads3mo agoHugging Face09yuyijiong /Multi-doc-QA-CommonCrawl Update on December 24, 2023: Improve the format of answers: force all answers to be provide the refered original text first. English multi document Q&A data created using RedPajamaCommonCrawl data as reference text In the Raw dataset, each sample contains one reference document, 199 irrelevant documents, and a Q-A pair based on the reference document. It can be used to train models to extract the target information from a large number of documents. After filtering, integrating, and… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Multi-doc-QA-CommonCrawl.text1K<n<10K9 likes59 downloads3y agoHugging Face10GeoGPT-Research-Project /GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset. This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl: id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.texttext-generation10M<n<100M1 likes38 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.