SinclairSchneider/tweets_sample_2026
Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus A large, deliberately untargeted sample of public posts from X/Twitter, collected via Nitter by sweeping a 65,689-term newspaper vocabulary rather than a topical keyword set. It is built as a background / reference corpus: a baseline of "what was being said in general" against which a topically targeted collection can be contrasted. It is the reference arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.
Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus
A large, deliberately untargeted sample of public posts from X/Twitter, collected via Nitter by sweeping a 65,689-term newspaper vocabulary rather than a topical keyword set.
It is built as a background / reference corpus: a baseline of "what was being said in general" against which a topically targeted collection can be contrasted. It is the reference arm of a narrative-detection study, not a curated dataset about any particular subject.
[!IMPORTANT] Two columns do not contain what their names say. `quotes` holds like counts and `likes` holds view counts. See Engagement columns are mislabeled before using any engagement metric.
At a glance
Loading
The corpus is 25.7 GB; stream it unless you have the RAM to spare.
from datasets import load_dataset
# streaming — no full download, constant memory
ds = load_dataset("SinclairSchneider/tweets_sample_2026", split="train", streaming=True)
print(next(iter(ds)))
# or materialize the whole thing (needs ~25 GB disk + a lot of RAM)
ds = load_dataset("SinclairSchneider/tweets_sample_2026", split="train")For analytical work, querying the Parquet files directly is far more efficient than going through datasets, because you can push column selection down to the file:
import duckdb
duckdb.sql("""
SELECT search_term, count(*) AS n
FROM 'hf://datasets/SinclairSchneider/tweets_sample_2026/data/*.parquet'
GROUP BY 1 ORDER BY n DESC LIMIT 20
""").show()The search vocabulary
search_terms.txt (shipped with this dataset) contains 65,689 lowercase, alphabetically sorted, single-token terms spanning aachen → übungsplatz. It is a general newspaper vocabulary — overwhelmingly ordinary words, not a political keyword list:
aachen, aal, aalglatt, aasgeier, abandon, abandoned, abarbeiten, abartig, abattoirs, abba, …
überziehend, überzieht, überzogen, üblich, übrig, übriggebliebener, übung, übungsplatzDespite the file name (political_reference_newspapers.txt), the list is not topically political; the "political" label refers to the study it was assembled for. It mixes German and English vocabulary, which is why the corpus skews English (see Language composition).
Hits are very unevenly distributed across terms:
The highest-yield terms are English and mostly generic: leftie (63,980), statistic (58,708), bootlicker (48,482), screenings (47,040), stadion (46,015), oversold (44,747), preferable (42,918).
Data fields
Known issues and caveats
Engagement columns are mislabeled
The scraper reads the four tweet-stat elements of a Nitter tweet card by position and assigns them the names comments, retweets, quotes, likes. On the instances actually used, positions 3 and 4 are likes and views — there is no quote count in the markup — so the last two columns are shifted by one:
The ratios make the shift unambiguous: the quotes column runs at 4.6× the retweet count, which is the canonical likes-to-retweets ratio (real quote counts are a fraction of retweets). The likes column runs at 99.7× retweets and 28.4× the `quotes` column, with a maximum of 1.54 billion — impossible for likes, entirely normal for views. Only 0.71 % of rows are zero, consistent with X displaying a view count on nearly every post.
Rename on load:
df = df.rename(columns={"comments": "replies", "quotes": "likes", "likes": "views"})A true quote count does not exist in this dataset.
search_term is lossy in two ways
- It is one arbitrary matching term, not all of them. A tweet containing several vocabulary words was returned by several queries, but deduplication kept only the first row encountered in filesystem-glob order.
search_termis therefore a term that matched, chosen arbitrarily — never treat it as the complete set of matching terms, and do not use it as a label or as a frequency estimate for that term in the corpus. - Long terms are truncated to their last 20 characters. The value is recovered from a filename built with
term[-20:], so 455 of the 65,689 terms (0.69 %) appear clipped —entscheidungskompetenz→tscheidungskompetenz,nachfolgeorganisation→achfolgeorganisation. Truncation can in principle also merge two distinct long terms that share a suffix.
Dates fall outside the requested window
Although the sweep requested since=2026-01-01, 144,215 rows (0.19 %) predate 2026, reaching back to 2006-03-24 — the search backend does not honour the date bound strictly. Filter on date yourself if you need a clean window.
The 2026 month-over-month growth is a collection artefact, not a signal about platform activity — scraping ran forward in time and later months had more instance capacity available. Do not read the monthly curve as a volume trend.
Language composition
No language filter was applied, and the vocabulary itself is mixed, so the corpus is majority English despite being assembled from a German newspaper vocabulary. A stopword heuristic over a 6,639,818-tweet sample:
This is a coarse heuristic, not a trained language ID. Run a proper detector (fastText lid.176, GlotLID) if language is load-bearing for your work.
Sampling bias
This is not a random sample of X, and no weighting will make it one:
- Coverage is bounded by the vocabulary. A tweet containing none of the 65,689 terms cannot appear — which biases against very short posts, emoji-only posts, other languages, and heavy slang.
- Nitter search returns what the search backend chooses to return, with its own opaque ranking and recall limits. Unlimited
--max_tweetsdoes not mean exhaustive retrieval. - Instance availability varied over the collection period, so coverage is uneven across time and across terms — 6,306 of the 65,689 vocabulary terms (9.6 %) yielded nothing at all, and it is not determinable after the fact whether a given term genuinely had no matches or was simply never served successfully.
- Engagement-related selection is unknown: high-visibility tweets are plausibly over-represented.
Other
- Content reflects the moment of scraping. Tweets later deleted, edited, or made private are still here; counts are frozen at scrape time and no longer match the platform.
- Retweets (4.75 %) duplicate the text of the original post — filter
is-retweetfor text work. dateis a string. Cast it before doing time arithmetic.
Intended use
Suited to: building background/reference distributions for narrative and framing detection, term-frequency baselines, corpus-linguistic study of platform language, retrieval and topic-model development, and pretraining or domain adaptation of social-media models.
Not suited to: measuring the prevalence of anything on X (the sampling frame forbids it), tracking trends over time (the monthly curve is an artefact), studying individuals or accounts, or any engagement analysis that has not first corrected the mislabeled columns.
