Omarrran/StackPulse_778K_QnA_Code_dataset
π» StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015β2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicatesβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.
π» StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015β2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.
π Files in This Dataset
ποΈ Schema Reference
π Dataset Statistics
Overview
- Total questions: 778,929
- Date range: 2015-02-14 β 2022-09-25
- Missing years: 2019, 2021 (sampling gaps)
- Unique tags: 41,754
- Zero nulls in all core columns (2 questions have empty tags)
Score Distribution
- Mean: 0.69 | Median: 0 | Std: 4.77
- Range: -27 to 1,061
- Negative score: 63,415 (8.14%)
- Zero score: 461,977 (59.31%) β majority never upvoted
- Score β₯ 10: 7,431 (0.95%)
- Score β₯ 100: 246 (0.032%)
View Count Distribution
- Mean: 794 | Median: 66 | P95: 2,599 | P99: 11,513
- Max: 915,870 views
- Popular (>P95): 38,944 (5.00%)
- Viral (>P99): 7,789 (1.00%)
Answer Count
- Unanswered: 224,733 (28.85%)
- 1 answer: 387,399 (49.73%)
- 2+ answers: 166,797 (21.42%)
- Max answers on a single question: 36
Question Body
- Has code block: 595,679 (76.47%)
- Avg code blocks per question: 1.56
- Avg body length: 557 chars | Median: 447 chars
- Avg title length: 59 chars
Tags
- Avg tags per question: 3.01
- 5 tags (SO max): 117,671 (15.1%)
- 1 tag: 92,409 (11.9%)
Questions Per Year
- -1: 2,694
- 2015: 99,664
- 2016: 120,887
- 2017: 20,791
- 2018: 20,096
- 2020: 99,676
- 2022: 415,121 (Note: 2017/2018 low counts reflect sampling focus; 2022 dominates at 53%)
Top 10 Tags
- python (98,154)
- javascript (87,156)
- java (53,052)
- c# (40,178)
- android (38,543)
- html (36,572)
- php (34,656)
- reactjs (31,727)
- css (24,770)
- r (21,694)
Unanswered Rate by Top Tags
- node.js: 36.46% unanswered (hardest to answer)
- reactjs: 35.31%
- flutter: 33.44%
- typescript: 31.19%
- android: 30.49%
- python: 28.69%
- java: 26.88%
- css: 19.59% (easiest to get answered)
- jquery: 18.30%
β οΈ Known Issues & Caveats
- YEAR GAPS: Years 2019 and 2021 are absent β this is a sampling artifact, not a gap in SO activity. Do not use for temporal trend analysis without noting this.
- 2022 DOMINANCE: 415,121 questions (53%) are from 2022. The dataset skews heavily toward recent questions. Stratify by year if balance matters.
- RAW HTML:
question_bodycontains raw HTML including<,>,<pre><code>blocks. Usebody_textfor NLP tasks. Usequestion_bodyfor HTML-aware or code-extraction tasks.
- SCORE SKEW: 59.3% of questions have score=0. Mean (0.69) is misleading. Use
score_bucketoris_highly_votedfor classification tasks.
- VIEW COUNT SKEW: Mean (794) is 12Γ the median (66) due to viral questions. Use log-transformed view_count for regression tasks.
- PIPE-SEPARATED TAGS: The
tagscolumn uses|as delimiter e.g.python|pandas|dataframe. Split withstr.split("|")before use.
- CODE PLACEHOLDER: In
body_text, all<pre>...</pre>blocks are replaced with the token[CODE]. The original HTML is preserved inquestion_body.
- DUPLICATE IDs: 2 exact duplicates were found and removed during processing.
π Quick Start
pandas
import pandas as pd
# Full dataset
df = pd.read_csv("data/stackoverflow_778k_full.csv")
# High quality only (score >= 5, answered)
hq = pd.read_csv("data/stackoverflow_high_quality.csv")
# Python questions only
py = pd.read_csv("data/stackoverflow_python.csv")
# Unanswered questions (good for difficulty modeling)
ua = pd.read_csv("data/stackoverflow_unanswered.csv")
# Split tags into list
df["tag_list"] = df["tags"].str.split("|")
# Filter by year
df_2022 = df[df["year"] == 2022]
# Log-transform view count for regression
import numpy as np
df["log_views"] = np.log1p(df["view_count"])HuggingFace datasets
from datasets import load_dataset
REPO = "Omarrran/StackPulse_778K_QnA_Code_dataset"
# Full 778K
ds = load_dataset(REPO, "full")
# High quality only
hq = load_dataset(REPO, "high_quality")
# Python questions
py = load_dataset(REPO, "python")
# Unanswered questions
ua = load_dataset(REPO, "unanswered")
# Convert to pandas
df = ds["train"].to_pandas()π¬ Suggested Research Tasks
π Citation
@dataset{malik2026stackoverflow,
author = {Malik, Omar Haq Nawaz},
title = {StackOverflow-778K: Multi-Year Developer Q&A Dataset},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset},
questions = {778929},
years = {2015-2022},
license = {Apache-2.0}
}π€ Author
Omar Haq Nawaz Malik (HuggingFace: Omarrran) AI Engineer & NLP Researcher | BITS Pilani | Srinagar, Kashmir
