datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackexchange-clustering
StackExchangeClustering.v2
An MTEB dataset
Massive Text Embedding Benchmark
Clustering of titles from 121 stackexchanges. Clustering of 25 sets, each with 10-50 classes, and each class with 100 - 1000 sentences.
Task category
t2c
Domains
Web, Written
Reference
https://arxiv.org/abs/2104.07081
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/stackexchange-clustering.stackexchange_filtered
Stack Exchange
Description
StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.stackexchange
Stack Exchange
Description
StackExchange is a collection of Q&A communities spanning a wide variety of topics.
While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.
Instead, each site can provide a logged in user with a custom url to download the dump for that site.
This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange.stackexchange-clustering-p2p
StackExchangeClusteringP2P.v2
An MTEB dataset
Massive Text Embedding Benchmark
Clustering of title+body from stackexchange. Clustering of 5 sets of 10k paragraphs and 5 sets of 5k paragraphs.
Task category
t2c
Domains
Web, Written
Reference
https://arxiv.org/abs/2104.07081
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["StackExchangeClusteringP2P.v2"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/stackexchange-clustering-p2p.physics-stackexchangestackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.stackexchange-markdown
Marin Markdownified StackExchange
Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training.
Value
Tokens
20 413 785 853
Primary source
https://archive.org/details/stackexchange
File format
JSONL
License
CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.stackexchange-all
StackExchange All Communities — Preprocessed
All 362 StackExchange communities (including Stack Overflow) processed into
QA pairs and standalone questions, ready for LLM pretraining and instruction tuning.
Stats
Field
Value
Communities
362
Dump date
March 2026
License
CC-BY-SA 4.0
Record Types
instruction — Question + Answer pair (qa_pair)
text — Unanswered question (standalone)
Format
{
"id": "uuid-v4"… See the full description on the dataset page: https://huggingface.co/datasets/raj2708/stackexchange-all.law-stack-exchange
Dataset Card for Law Stack Exchange Dataset
Dataset Summary
Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation".
Citation Information
@inproceedings{li-etal-2022-parameter,
title = "Parameter-Efficient Legal Domain Adaptation",
author = "Li, Jonathan and
Bhambhoria, Rohan and
Zhu, Xiaodan",
booktitle = "Proceedings of the Natural Legal Language Processing Workshop 2022",
month = dec… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/law-stack-exchange.Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.MedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API.
robotics_stackexchange_com_QLORAstackexchange-attrpromptThis is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias.
Checkout the paper: https://arxiv.org/abs/2306.15895 for details.
label.txt: the label name for each class
train.jsonl: The original training set.
valid.jsonl: The original validation set.
test.jsonl: The original test set.
simprompt.jsonl: The training data generated by the simple prompt.
attrprompt.jsonl: The training data generated by the attributed prompt.
arena_emb_stackexchangeEmbeddings used for the MTEB Arena.
You can download this repo via
git clone https://hf.co/datasets/mteb/arena_emb_stackexchange
cd arena_emb_stackexchange
git lfs pull
or just download individual files, e.g. wget https://hf.co/datasets/mteb/arena_emb_stackexchange/resolve/main/emb_stackexchange_GritLM__GritLM-7B.json.aa
As there is an upload limit of 50GB per file, we have split files using e.g. split --number=l/6 emb_stackexchange_GritLM__GritLM-7B.json… See the full description on the dataset page: https://huggingface.co/datasets/mteb-arena/arena_emb_stackexchange.tigerbot-stackexchange-qa-en-0.5mTigerbot 基于stackexchange问答站点dump数据生成sft数据集
原始来源:https://archive.org/details/stackexchange
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-stackexchange-qa-en-0.5m')
reverseengineering.stackexchange.com3dprinting.stackexchange.comarena-stackexchange
Dataset used for Stackexchange in MTEB/arena
Overview
The mteb/arena-stackexchange dataset is a curated collection of Stack Exchange questions and answers, designed for use in the MTEB (Massive Text Embedding Benchmark) Arena. This dataset allows various embedding models to compete and be ranked based on their performance on Stack Exchange content.
What is Stack Exchange?
Stack Exchange is a network of question-and-answer (Q&A) websites on topics in diverse… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arena-stackexchange.preference-data-math-stack-exchangeThe preference dataset is derived from the stack exchange dataset which contains questions and answers from the Stack Overflow Data Dump. This contains questions and answers for various topics. For this work, we used only question and answers from math.stackexchange.com sub-folder.
The questions are grouped with answers that are assigned a score corresponding to the Anthropic paper:
score = log2 (1 + upvotes) rounded to the nearest integer, plus 1 if the answer was accepted by the questioner… See the full description on the dataset page: https://huggingface.co/datasets/prhegde/preference-data-math-stack-exchange.3dprinting.stackexchange.compile_stackexchangeredpajama-pile-stackexchange-refined-by-data-juicer
RedPajama & The Pile -- StackExchange (refined by Data-Juicer)
A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB).
Dataset Information
Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.stackexchangeelectronics.stackexchange.comiot.stackexchange.comstackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
robotics.stackexchange.comstackexchange-200-wordsiot.meta.stackexchange.comstackexchange_dolma
