datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackexchange_filtered
Stack Exchange
Description
StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.stackexchange
Stack Exchange
Description
StackExchange is a collection of Q&A communities spanning a wide variety of topics.
While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.
Instead, each site can provide a logged in user with a custom url to download the dump for that site.
This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange.stack-exchange-paired
StackExchange Paired
This is a processed version of the HuggingFaceH4/stack-exchange-preferences. The following steps were applied:
Parse HTML to Markdown with markdownify
Create pairs (response_j, response_k) where j was rated better than k
Sample at most 10 pairs per question
Shuffle the dataset globally
This dataset is designed to be used for preference learning. The processing notebook is in the repository as well.
stackexchange-markdown
Marin Markdownified StackExchange
Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training.
Value
Tokens
20 413 785 853
Primary source
https://archive.org/details/stackexchange
File format
JSONL
License
CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.smollm3-stackexchange
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
This repository contains pre-processed datasets used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior.
The datasets include pre-pretraining (PPT) data and pre-training (PT) data mixtures used across multiple model scales and configurations. For full details on dataset composition, preprocessing, and training, please refer to the GitHub repository.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stackexchange.the_pile_stack_exchangeThis dataset is part of EleutherAI/The Pile dataset and is a dataset for Language Models from processing stackexchange data dump, which is an anonymized dump of all user-contributed content on the Stack Exchange network.stackexchange-space-qa
Stack Exchange Space Q&A
Credit: NASA/DOE/Fermi LAT Collaboration
Part of a dataset collection on Hugging Face.
Dataset description
This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.japanese-stackexchange
japanese-stackexchange
英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
日本語翻訳された StackExchange ではないです。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.qualc-stackexchange-en
QualC Stack Exchange English
A cleaned and validated English Stack Exchange corpus prepared for large language model (LLM) pretraining.
This dataset is part of the QualC Foundation Corpus, an open collection of high-quality datasets intended for training multilingual foundation models.
Dataset Information
Language: English
Records: 999,832
Format: Hugging Face Dataset
Schema: Flat Text
License: CC BY-SA 4.0 (inherits the original Stack Exchange content license)… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-stackexchange-en.redpajama-pile-stackexchange-refined-by-data-juicer
RedPajama & The Pile -- StackExchange (refined by Data-Juicer)
A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB).
Dataset Information
Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.stackexchange-duplicates_marathi
Stackexchange-Duplicates Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Stackexchange-Duplicates Marathi dataset is a meticulously curated collection of 304525 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/stackexchange-duplicates_marathi.stack-exchange-paired-mini
StackExchange Paired Mini (100 samples)
This is a subset of the StackExchange Paired lvwerra/stack-exchange-paired dataset.
Disclaimer
For licensing or any other related detail, please refer to the original dataset linked above.
