rustemgareev/tg-contest
Telegram contest corpus 790,166 news articles collected from Telegram channels, 27 April – 17 May 2020. There are no labels: the corpus is unlabeled on purpose — annotation is done by the contest participants. Built for content-classification work: zero-shot and few-shot annotation by participants, training on your own labels, filtering by outlet or date, and information extraction. What is inside The corpus was collected from Telegram posts linking to articles.… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/tg-contest.
Telegram contest corpus
790,166 news articles collected from Telegram channels, 27 April – 17 May 2020. There are no labels: the corpus is unlabeled on purpose — annotation is done by the contest participants.
Built for content-classification work: zero-shot and few-shot annotation by participants, training on your own labels, filtering by outlet or date, and information extraction.
What is inside
The corpus was collected from Telegram posts linking to articles. The upstream HTML files were deleted by the converter, so the .md files are the only source left. Therefore text is stored verbatim, with all markdown markup intact.
Columns
About text
The body is stored verbatim, with two normalizations, both reversible:
CRLF → LF(recorded in thecrlfflag),- leading and trailing blank lines stripped (pure visual noise).
Line breaks inside paragraphs and all markup (#, **, >, tables, links, images) are exactly as in the original. The first lines almost always contain a duplicated # title and a *date · author* byline — they are already available as separate columns, and tools/load_example.py shows how to strip them before training.
Loading
from datasets import load_dataset
ds = load_dataset("rustemgareev/tg-contest", split="train") # full
ds = load_dataset("rustemgareev/tg-contest", split="train", streaming=True) # streamed
# training text: body without the duplicated header
import re
def prep(row):
text = re.sub(r"^#\s+.*\n+", "", row["text"], count=1)
return re.sub(r"\n?\*[^*\n]{0,120}\*\n", "\n", text, count=1).strip()
# attach participant annotations
import pandas as pd
annot = pd.read_csv("annotations.csv")
labels = dict(zip(annot["id"], annot["label"]))
ds = ds.filter(lambda r: r["id"] in labels).map(lambda r: {**r, "label": labels[r["id"]]})Files live at data/train-00000-of-00021-20200427.parquet … — they can also be read directly with pyarrow.parquet, without datasets.
Known limitations
- No labels. The corpus is unlabeled by design — participants annotate it. There is no
labelfield in this dataset. - `lang` is unreliable. It is a heuristic from the upstream converter, not ground truth. The converter also discarded everything it did not recognize as
ruoren, so only those two languages remain and the size of the discarded portion is unknown. If language matters, re-detect it fromtext. - The source channel is not in the corpus. The Telegram channel a post came from was never recorded: the filename is a bare
message_id, andsiteis the article's outlet, not the channel. Grouping "by channel" is therefore impossible; grouping "by outlet" works. - 234 rows have invalid YAML in their front matter: the converter escaped only
\and", but not line breaks. Values were recovered with a line-based fallback parser and are flaggedyaml_valid=False. - 46 rows contain U+FFFD — a replacement character left over from conversion (
utf8_ok=True; the replacement was already in the source). - 1,051 rows duplicate a `message_id` with different content; no two files are byte-identical. Check
id_repeatwhen building splits. - Duplicates were not removed. They amount to 0.43 %, so removal would gain almost nothing, while 5.4 % of normalized-title matches are mostly different articles with similar headlines and would have been discarded wrongly. Instead of deletion there are the
dup_groupandsha1_bodycolumns — the decision is yours.
Integrity
sha256sum -c stats/shards.sha256 # Linux
Get-FileHash data\*.parquet -Algorithm SHA256 # Windowsstats/manifest.json holds per-day statistics: row counts, language distribution, broken front matter, missing fields, shard size. stats/duplicates.json and stats/duplicates.csv describe the duplicate groups.
Packing was verified by an end-to-end round-trip: 400 randomly sampled rows were re-read from the original .md files and compared on sha1_raw and every field — byte-identical, 400/400.
How this was built
Source: 790,166 files under YYYYMMDD/HH/<message_id>.md (21 days × 24 hours), 3.63 GB, CRLF, UTF-8, front matter present in every file.
build_parquet.py— a single pass: a pool of 16 threads doingread()only, parsing in the main process. The source sits on Storage Spaces backed by a single spinning HDD (~4 ms per file), so extra read threads do not help, while threads fighting over the GIL inside the regexes actively hurt. Writes one shard per day, then an atomic rename.assign_dup_groups.py— labels duplicates using the finished shards, not a second pass over the HDD.validate_parquet.py— schema, checksums, round-trip against the source files, anddup_groupconsistency.
Full pass over the corpus: 37 minutes; the whole pipeline including the second stage takes about 45 minutes. The scripts live in the project's git repository and are not published as part of the dataset.
