datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-exchange-preferences
Dataset Card for H4 Stack Exchange Preferences Dataset
Dataset Summary
This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training.
Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers.
This data could also be used for instruction fine-tuning and language model training.
The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.stack-exchange-paired
StackExchange Paired
This is a processed version of the HuggingFaceH4/stack-exchange-preferences. The following steps were applied:
Parse HTML to Markdown with markdownify
Create pairs (response_j, response_k) where j was rated better than k
Sample at most 10 pairs per question
Shuffle the dataset globally
This dataset is designed to be used for preference learning. The processing notebook is in the repository as well.
stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.math_stackexchange_qa
Math StackExchange Curated (Parquet, CC BY-SA 4.0)
This dataset is a curated collection of Math StackExchange (MSE) Q&A pairs packaged in Parquet format.Each sample contains a problem (title, question_body), its corresponding answer (answer_body), the original MSE tag string (tags), and a flag indicating whether the answer was accepted (accepted).
This dataset includes content derived from the Math StackExchange public data dump (CC BY-SA 4.0, © Stack Exchange Inc.).This derived… See the full description on the dataset page: https://huggingface.co/datasets/glopezas/math_stackexchange_qa.MedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API.
stackexchangeThis dataset is based entirely on HuggingFaceH4/stack-exchange-preferences, but it has been restructured.
All HTML tags have been cleaned out, and the answers column has been turned into the answer column, so instead of answers being stored in JSON format there is now a row for each answer.
Furthermore there is a separate file for every forum instead of a single file.
stack-exchange-dataset
Overview
This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv.
Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021.
cs.tsv --> Computer Science SE
ds.csv --> Data Science SE
p.csv --> Political Science SE
File Structure
Each file has the following columns:
id: the question id
title: the title of the question
body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.stackexchange-space-qa
Stack Exchange Space Q&A
Credit: NASA/DOE/Fermi LAT Collaboration
Part of a dataset collection on Hugging Face.
Dataset description
This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.japanese-stackexchange
japanese-stackexchange
英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
日本語翻訳された StackExchange ではないです。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.ukrainian-stackexchange
Ukrainian StackExchange Dataset
This repository contains a dataset collected from the Ukrainian StackExchange website.
The parsed date is 02/04/2023.
The dataset is in JSON format and includes text data parsed from the website https://ukrainian.stackexchange.com/.
Dataset Description
The Ukrainian StackExchange Dataset is a rich source of text data for tasks related to natural language processing, machine learning, and data mining in the Ukrainian language. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/zeusfsx/ukrainian-stackexchange.qualc-stackexchange-en
QualC Stack Exchange English
A cleaned and validated English Stack Exchange corpus prepared for large language model (LLM) pretraining.
This dataset is part of the QualC Foundation Corpus, an open collection of high-quality datasets intended for training multilingual foundation models.
Dataset Information
Language: English
Records: 999,832
Format: Hugging Face Dataset
Schema: Flat Text
License: CC BY-SA 4.0 (inherits the original Stack Exchange content license)… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-stackexchange-en.stackexchange-duplicates_marathi
Stackexchange-Duplicates Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Stackexchange-Duplicates Marathi dataset is a meticulously curated collection of 304525 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/stackexchange-duplicates_marathi.stack-exchange-paired-smallstack-exchange-paired-mini
StackExchange Paired Mini (100 samples)
This is a subset of the StackExchange Paired lvwerra/stack-exchange-paired dataset.
Disclaimer
For licensing or any other related detail, please refer to the original dataset linked above.
Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/LiuXH648/Law-StackExchange.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.
