Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /stackexchange_filtered Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.texttext-generation10M<n<100M10 likes4.2k downloads1y agoHugging Face02common-pile /stackexchange Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics. While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive. Instead, each site can provide a logged in user with a custom url to download the dump for that site. This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange.texttext-generation10M<n<100M8 likes4k downloads1y agoHugging Face03marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes752 downloads1y agoHugging Face04juliensimon /stackexchange-space-qa Stack Exchange Space Q&A Credit: NASA/DOE/Fermi LAT Collaboration Part of a dataset collection on Hugging Face. Dataset description This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.tabularquestion-answering10K<n<100K0 likes96 downloads20d agoHugging Face05p1atdev /japanese-stackexchange japanese-stackexchange 英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 日本語翻訳された StackExchange ではないです。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.tabulartext-generation10K<n<100K3 likes74 downloads3y agoHugging Face06nandhakumarms /qualc-stackexchange-en QualC Stack Exchange English A cleaned and validated English Stack Exchange corpus prepared for large language model (LLM) pretraining. This dataset is part of the QualC Foundation Corpus, an open collection of high-quality datasets intended for training multilingual foundation models. Dataset Information Language: English Records: 999,832 Format: Hugging Face Dataset Schema: Flat Text License: CC BY-SA 4.0 (inherits the original Stack Exchange content license)… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-stackexchange-en.texttext-generation100K<n<1M0 likes40 downloads2mo agoHugging Face07datajuicer /redpajama-pile-stackexchange-refined-by-data-juicer RedPajama & The Pile -- StackExchange (refined by Data-Juicer) A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB). Dataset Information Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.texttext-generationn<1K0 likes36 downloads3y agoHugging Face08Singhchandann /stackexchange-duplicates_marathigated Stackexchange-Duplicates Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Stackexchange-Duplicates Marathi dataset is a meticulously curated collection of 304525 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/stackexchange-duplicates_marathi.texttext-classification100K<n<1M0 likes17 downloads1y agoHugging Face09alvarobartt /stack-exchange-paired-mini StackExchange Paired Mini (100 samples) This is a subset of the StackExchange Paired lvwerra/stack-exchange-paired dataset. Disclaimer For licensing or any other related detail, please refer to the original dataset linked above. texttext-generationn<1K0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.