Team Ai
Datasetpublic

Omarrran/StackPulse_778K_QnA_Code_dataset

πŸ’» StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes182downloads
Dataset Card

πŸ’» StackOverflow-778K: Multi-Year Developer Q&A Dataset

Dataset Summary

A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use.

Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.


πŸ“ Files in This Dataset

FileFormatRowsDescription
stackoverflow778kfull.csvCSV778,929Complete dataset
stackoverflow_unanswered.csvCSV224,733Questions with 0 answers
stackoverflowwithcode.csvCSV595,679Questions containing code blocks
stackoverflowhighquality.csvCSV20,205Score β‰₯ 5 and answered
stackoverflow_python.csvCSV107,083Python-tagged questions
stackoverflow_javascript.csvCSV87,367JavaScript-tagged questions
stackoverflow_java.csvCSV54,077Java-tagged questions
stackoverflow_csharp.csvCSV40,420C#-tagged questions
stackoverflow_android.csvCSV42,004Android-tagged questions
*.jsonl versionsJSONLsameHuggingFace-native format for all above

πŸ—οΈ Schema Reference

ColumnTypeDescription
idint64Unique Stack Overflow question ID
titlestringQuestion title
question_bodystringRaw HTML body (includes <pre><code> blocks)
body_textstringPlain text body (HTML stripped, code replaced with [CODE])
tagsstringPipe-separated tags e.g. `python\pandas\dataframe`
scoreint64Net upvotes (can be negative)
creation_datestringISO 8601 UTC creation timestamp
yearint32Year extracted from creation_date
monthint32Month (1–12)
hourint32Hour of day (0–23, UTC)
dayofweekint32Day of week (0=Monday, 6=Sunday)
view_countint64Total question views
answer_countint64Number of answers received
body_lenint64Plain text body character length
title_lenint64Title character length
tag_countint64Number of tags (1–5)
codeblockcntint64Number of <pre> code blocks in body
has_codeboolTrue if question contains at least one code block
is_unansweredboolTrue if answer_count == 0
is_popularboolTrue if view_count > 95th percentile (~2,599)
is_viralboolTrue if view_count > 99th percentile (~11,513)
ishighlyvotedboolTrue if score >= 10
is_negativeboolTrue if score < 0
score_bucketstring"negative" / "zero" / "low" / "medium" / "high"

πŸ“ˆ Dataset Statistics

Overview

  • β€”Total questions: 778,929
  • β€”Date range: 2015-02-14 β†’ 2022-09-25
  • β€”Missing years: 2019, 2021 (sampling gaps)
  • β€”Unique tags: 41,754
  • β€”Zero nulls in all core columns (2 questions have empty tags)

Score Distribution

  • β€”Mean: 0.69 | Median: 0 | Std: 4.77
  • β€”Range: -27 to 1,061
  • β€”Negative score: 63,415 (8.14%)
  • β€”Zero score: 461,977 (59.31%) β€” majority never upvoted
  • β€”Score β‰₯ 10: 7,431 (0.95%)
  • β€”Score β‰₯ 100: 246 (0.032%)

View Count Distribution

  • β€”Mean: 794 | Median: 66 | P95: 2,599 | P99: 11,513
  • β€”Max: 915,870 views
  • β€”Popular (>P95): 38,944 (5.00%)
  • β€”Viral (>P99): 7,789 (1.00%)

Answer Count

  • β€”Unanswered: 224,733 (28.85%)
  • β€”1 answer: 387,399 (49.73%)
  • β€”2+ answers: 166,797 (21.42%)
  • β€”Max answers on a single question: 36

Question Body

  • β€”Has code block: 595,679 (76.47%)
  • β€”Avg code blocks per question: 1.56
  • β€”Avg body length: 557 chars | Median: 447 chars
  • β€”Avg title length: 59 chars

Tags

  • β€”Avg tags per question: 3.01
  • β€”5 tags (SO max): 117,671 (15.1%)
  • β€”1 tag: 92,409 (11.9%)

Questions Per Year

  • β€”-1: 2,694
  • β€”2015: 99,664
  • β€”2016: 120,887
  • β€”2017: 20,791
  • β€”2018: 20,096
  • β€”2020: 99,676
  • β€”2022: 415,121 (Note: 2017/2018 low counts reflect sampling focus; 2022 dominates at 53%)

Top 10 Tags

  • β€”python (98,154)
  • β€”javascript (87,156)
  • β€”java (53,052)
  • β€”c# (40,178)
  • β€”android (38,543)
  • β€”html (36,572)
  • β€”php (34,656)
  • β€”reactjs (31,727)
  • β€”css (24,770)
  • β€”r (21,694)

Unanswered Rate by Top Tags

  • β€”node.js: 36.46% unanswered (hardest to answer)
  • β€”reactjs: 35.31%
  • β€”flutter: 33.44%
  • β€”typescript: 31.19%
  • β€”android: 30.49%
  • β€”python: 28.69%
  • β€”java: 26.88%
  • β€”css: 19.59% (easiest to get answered)
  • β€”jquery: 18.30%

⚠️ Known Issues & Caveats

  1. 1.YEAR GAPS: Years 2019 and 2021 are absent β€” this is a sampling artifact, not a gap in SO activity. Do not use for temporal trend analysis without noting this.
  1. 1.2022 DOMINANCE: 415,121 questions (53%) are from 2022. The dataset skews heavily toward recent questions. Stratify by year if balance matters.
  1. 1.RAW HTML: question_body contains raw HTML including &lt;, &gt;, <pre><code> blocks. Use body_text for NLP tasks. Use question_body for HTML-aware or code-extraction tasks.
  1. 1.SCORE SKEW: 59.3% of questions have score=0. Mean (0.69) is misleading. Use score_bucket or is_highly_voted for classification tasks.
  1. 1.VIEW COUNT SKEW: Mean (794) is 12Γ— the median (66) due to viral questions. Use log-transformed view_count for regression tasks.
  1. 1.PIPE-SEPARATED TAGS: The tags column uses | as delimiter e.g. python|pandas|dataframe. Split with str.split("|") before use.
  1. 1.CODE PLACEHOLDER: In body_text, all <pre>...</pre> blocks are replaced with the token [CODE]. The original HTML is preserved in question_body.
  1. 1.DUPLICATE IDs: 2 exact duplicates were found and removed during processing.

πŸš€ Quick Start

pandas

python
import pandas as pd

# Full dataset
df = pd.read_csv("data/stackoverflow_778k_full.csv")

# High quality only (score >= 5, answered)
hq = pd.read_csv("data/stackoverflow_high_quality.csv")

# Python questions only
py = pd.read_csv("data/stackoverflow_python.csv")

# Unanswered questions (good for difficulty modeling)
ua = pd.read_csv("data/stackoverflow_unanswered.csv")

# Split tags into list
df["tag_list"] = df["tags"].str.split("|")

# Filter by year
df_2022 = df[df["year"] == 2022]

# Log-transform view count for regression
import numpy as np
df["log_views"] = np.log1p(df["view_count"])

HuggingFace datasets

python
from datasets import load_dataset

REPO = "Omarrran/StackPulse_778K_QnA_Code_dataset"

# Full 778K
ds = load_dataset(REPO, "full")

# High quality only
hq = load_dataset(REPO, "high_quality")

# Python questions
py = load_dataset(REPO, "python")

# Unanswered questions
ua = load_dataset(REPO, "unanswered")

# Convert to pandas
df = ds["train"].to_pandas()

πŸ”¬ Suggested Research Tasks

TaskConfigKey Columns
Answer prediction (binary)fulltitle, bodytext, tags β†’ isunanswered
Score regressionfulltitle, body_text, tags β†’ score
View count predictionfulltitle, tags, score β†’ log(view_count)
Tag recommendationfulltitle, body_text β†’ tags
Code vs no-code classificationfullbodytext β†’ hascode
Question quality scoringfulltitle, bodytext β†’ scorebucket
LLM fine-tuning (Q&A)high_qualitytitle + body_text as prompt
Difficulty estimationfulltags β†’ unanswered rate per tag
Time-of-day analysisfullhour, dayofweek β†’ view_count / score
Language-specific modelingpython/javascript/javaany

πŸ“‹ Citation

bibtex
@dataset{malik2026stackoverflow,
  author    = {Malik, Omar Haq Nawaz},
  title     = {StackOverflow-778K: Multi-Year Developer Q&A Dataset},
  year      = {2026},
  publisher = {HuggingFace},
  url       = {https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset},
  questions = {778929},
  years     = {2015-2022},
  license   = {Apache-2.0}
}

πŸ‘€ Author

Omar Haq Nawaz Malik (HuggingFace: Omarrran) AI Engineer & NLP Researcher | BITS Pilani | Srinagar, Kashmir