datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.stackoverflow-clean
Dataset Card for "stackoverflow-clean"
More Information needed
stackoverflowstackoverflow_python
Dataset Card for "stackoverflow_python"
Dataset Summary
This dataset comes originally from kaggle.
It was originally split into three tables (CSV files) (Questions, Answers, and Tags)
now merged into a single table. Each row corresponds to a pair (question-answer) and
their associated tags.
The dataset contains all questions asked between August 2, 2008 and Ocotober 19, 2016.
Supported Tasks and Leaderboards
This might be useful for open-domain… See the full description on the dataset page: https://huggingface.co/datasets/koutch/stackoverflow_python.stackoverflow_Llama-3.1-8B-Instruct_vocab_2000_laststackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.stackoverflow-unified-text-open-status-classification
Dataset Card for "stackoverflow-unified-text-open-status-classification"
More Information needed
stackoverflow-datasetstackoverflow_DL-related_questionsstackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
ja-stackoverflow
ja-stackoverflow
日本語版 Stack Overflow の スタック・オーバーフロー のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット
default サブセットから、 question と answers… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/ja-stackoverflow.stackoverflowVQA
Dataset Card for "stackoverflowVQA"
More Information needed
stack-overflowdumpstackoverflow_qa_python_Preprocessedstackoverflow_ERNIE-4.5-0.3B-PT_vocab_4000_laststackoverflow-unified-text-open-status-classification-sample
Dataset Card for "stackoverflow-open-status-classification"
More Information needed
stackoverflow_ERNIE-4.5-0.3B-PT_vocab_2000_laststackoverflow_ERNIE-4.5-0.3B-PT_vocab_1000_laststackoverflow_Llama-3.2-1B-Instruct_vocab_2000_laststackoverflow_ERNIE-4.5-0.3B-PT_vocab_500_laststackoverflow_q_and_a_sample
Description
GitHub repository: https://github.com/EshanJayasundara/Stackoverflow-Python-Q-and-A-Extractor.
GitHub repository contains the automated workflow for extracting the question and answer pairs from Stackoverflow.
This dataset contains the question-answer pairs extracted from Stackoverflow using Stack Exchange API v2.3 and used following endpoints,
/answers/{ids} GET
/questions GET
From 2020 January 1 to Today
1. Dataset description,
Contains only python… See the full description on the dataset page: https://huggingface.co/datasets/eshangj/stackoverflow_q_and_a_sample.stackoverflow-scraper
StackOverflow Scraper
Scrape Stack Overflow questions, answers, tags and user profiles through the public Stack Exchange API. Filter by tag, score, date, accepted status and full-text search. No login, no browser.
Rows in this dataset
16,719
Fields
47
Collector runs behind it
57
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/stackoverflow-scraper/ — 9,841 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/stackoverflow-scraper.stack-overflowstack-overflow-datasetstackoverflow_QAs
StackOverflow Q&A Dataset for Various Projects
Description
This dataset consists of Q&A data extracted from StackOverflow, related to different projects of CNCF (Cloud Native Computing Foundation) landscape. It includes the following three columns:
Question: The question asked on StackOverflow.
Answer: The corresponding answer to the question.
Tag: The name of the project to which the question and answer are related.
The data was collected using the Git Exchange API to… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/stackoverflow_QAs.stackoverflow-qa-top-300kstack-overflow
Stack Overflow Dataset
This dataset contains badge awards earned by users on Stack Overflow between January 1, 2022, and December 31, 2023. It includes 3,336 sequences with 187,836 events and 25 badge types, derived from the Stack Exchange Data Dump under the CC BY-SA 4.0 license. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
Update (2025-10-28): Added three timestamp fields (timestamp_event… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/stack-overflow.stack-overflow-description
Stack Overflow Description Dataset
This dataset contains badge awards earned by users on Stack Overflow between January 1, 2022, and December 31, 2023. It includes 3,336 sequences with 187,836 events and 25 badge types, derived from the Stack Exchange Data Dump under the CC BY-SA 4.0 license. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/stack-overflow-description.cognitive-traces-stackoverflow
Cognitive Traces — Stack Overflow
Dataset Description
This dataset contains cognitive trace annotations for the Stack Overflow dataset, produced by the multi-agent annotation framework described in:
Beyond the Click: A Framework for Inferring Cognitive Traces in Search
Saber Zerhoudi, Michael Granitzer. ECIR 2026.
Each user event (question, answer, comment, edit, vote) is annotated with a cognitive label from Information Foraging Theory (IFT), along with the full… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/cognitive-traces-stackoverflow.StackOverflow-TP4-1M
Dataset Details
Dataset Description
TP4 is a comprehensive dataset containing a curated collection of questions and answers from Stack Overflow. Focused on the realms of Python programming, NumPy, Pandas, TensorFlow, and PyTorch, TP4 includes essential attributes such as question ID, title, question body, answer body, associated tags, and score. This dataset is designed to facilitate research, analysis, and exploration of inquiries and solutions within the Python and… See the full description on the dataset page: https://huggingface.co/datasets/Syed-Hasan-8503/StackOverflow-TP4-1M.
