Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /stack-exchange-preferences Dataset Card for H4 Stack Exchange Preferences Dataset Dataset Summary This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training. Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers. This data could also be used for instruction fine-tuning and language model training. The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.textquestion-answering10M<n<100M136 likes6.8k downloads4y agoHugging Face02lvwerra /stack-exchange-paired StackExchange Paired This is a processed version of the HuggingFaceH4/stack-exchange-preferences. The following steps were applied: Parse HTML to Markdown with markdownify Create pairs (response_j, response_k) where j was rated better than k Sample at most 10 pairs per question Shuffle the dataset globally This dataset is designed to be used for preference learning. The processing notebook is in the repository as well. text-generation10M<n<100M150 likes1.7k downloads4y agoHugging Face03flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes1.2k downloads4y agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M11 likes1.2k downloads4y agoHugging Face05flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M13 likes783 downloads4y agoHugging Face06flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes531 downloads4y agoHugging Face07ymoslem /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.tabularquestion-answering10K<n<100K32 likes148 downloads1y agoHugging Face08glopezas /math_stackexchange_qa Math StackExchange Curated (Parquet, CC BY-SA 4.0) This dataset is a curated collection of Math StackExchange (MSE) Q&A pairs packaged in Parquet format.Each sample contains a problem (title, question_body), its corresponding answer (answer_body), the original MSE tag string (tags), and a flag indicating whether the answer was accepted (accepted). This dataset includes content derived from the Math StackExchange public data dump (CC BY-SA 4.0, © Stack Exchange Inc.).This derived… See the full description on the dataset page: https://huggingface.co/datasets/glopezas/math_stackexchange_qa.textquestion-answering1M<n<10M0 likes137 downloads11mo agoHugging Face09ymoslem /MedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. tabularquestion-answering1K<n<10K11 likes111 downloads3y agoHugging Face10kispeterzsm-szte /stackexchangeThis dataset is based entirely on HuggingFaceH4/stack-exchange-preferences, but it has been restructured. All HTML tags have been cleaned out, and the answers column has been turned into the answer column, so instead of answers being stored in JSON format there is now a row for each answer. Furthermore there is a separate file for every forum instead of a single file. textquestion-answering10M<n<100M2 likes101 downloads1y agoHugging Face11habedi /stack-exchange-dataset Overview This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv. Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021. cs.tsv --> Computer Science SE ds.csv --> Data Science SE p.csv --> Political Science SE File Structure Each file has the following columns: id: the question id title: the title of the question body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.tabulartext-classification10K<n<100K11 likes97 downloads8mo agoHugging Face12juliensimon /stackexchange-space-qa Stack Exchange Space Q&A Credit: NASA/DOE/Fermi LAT Collaboration Part of a dataset collection on Hugging Face. Dataset description This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.tabularquestion-answering10K<n<100K0 likes96 downloads20d agoHugging Face13p1atdev /japanese-stackexchange japanese-stackexchange 英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 日本語翻訳された StackExchange ではないです。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.tabulartext-generation10K<n<100K3 likes74 downloads3y agoHugging Face14zeusfsx /ukrainian-stackexchange Ukrainian StackExchange Dataset This repository contains a dataset collected from the Ukrainian StackExchange website. The parsed date is 02/04/2023. The dataset is in JSON format and includes text data parsed from the website https://ukrainian.stackexchange.com/. Dataset Description The Ukrainian StackExchange Dataset is a rich source of text data for tasks related to natural language processing, machine learning, and data mining in the Ukrainian language. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/zeusfsx/ukrainian-stackexchange.question-answering1K<n<10K5 likes51 downloads4y agoHugging Face15nandhakumarms /qualc-stackexchange-en QualC Stack Exchange English A cleaned and validated English Stack Exchange corpus prepared for large language model (LLM) pretraining. This dataset is part of the QualC Foundation Corpus, an open collection of high-quality datasets intended for training multilingual foundation models. Dataset Information Language: English Records: 999,832 Format: Hugging Face Dataset Schema: Flat Text License: CC BY-SA 4.0 (inherits the original Stack Exchange content license)… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-stackexchange-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face16Singhchandann /stackexchange-duplicates_marathigated Stackexchange-Duplicates Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Stackexchange-Duplicates Marathi dataset is a meticulously curated collection of 304525 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/stackexchange-duplicates_marathi.texttext-classification100K<n<1M0 likes17 downloads1y agoHugging Face17john02171574 /stack-exchange-paired-smallquestion-answering0 likes17 downloads1y agoHugging Face18alvarobartt /stack-exchange-paired-mini StackExchange Paired Mini (100 samples) This is a subset of the StackExchange Paired lvwerra/stack-exchange-paired dataset. Disclaimer For licensing or any other related detail, please refer to the original dataset linked above. texttext-generationn<1K0 likes12 downloads3y agoHugging Face19LiuXH648 /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/LiuXH648/Law-StackExchange.tabularquestion-answering10K<n<100K0 likes6 downloads9mo agoHugging Face20uvira007 /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.question-answering0 likes6 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.