Team Ai
Datasetpublic

rustemgareev/tg-contest

Telegram contest corpus 790,166 news articles collected from Telegram channels, 27 April – 17 May 2020. There are no labels: the corpus is unlabeled on purpose — annotation is done by the contest participants. Built for content-classification work: zero-shot and few-shot annotation by participants, training on your own labels, filtering by outlet or date, and information extraction. What is inside The corpus was collected from Telegram posts linking to articles.… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/tg-contest.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes66downloads
Dataset Card

Telegram contest corpus

790,166 news articles collected from Telegram channels, 27 April – 17 May 2020. There are no labels: the corpus is unlabeled on purpose — annotation is done by the contest participants.

Built for content-classification work: zero-shot and few-shot annotation by participants, training on your own labels, filtering by outlet or date, and information extraction.

What is inside

The corpus was collected from Telegram posts linking to articles. The upstream HTML files were deleted by the converter, so the .md files are the only source left. Therefore text is stored verbatim, with all markdown markup intact.

rows790,166
shards21 (one per day)
size1.08 GB (from 3.63 GB of source .md, zstd-9 ≈ 3.4×)
languageru — 444,739 (56.3 %), en — 345,427 (43.7 %)
mean text length3,064 characters, max 255,854
period2020-04-27 … 2020-05-17
duplicates3,396 rows (0.43 %)
labelsnone

Columns

columntypedescription
idint64Telegram message_id taken from the filename. Stable row key — this is what annotations are joined on
pathstringoriginal path YYYYMMDD/HH/<id>.md
daystringday folder, YYYYMMDD
hourint16hour folder, 00…23 (UTC)
datestringarticle:published_time from the source, ISO-8601, verbatim. May differ from the day folder
langstringru / en. Unreliable, see limitations below
titlestringheadline from the front matter
descriptionstringlead from the front matter, empty in 0.5 % of rows
urlstringlink to the original article
sitestringoutlet (og:site_name)
authorstringauthor, empty in 41 % of rows
author_urlstringauthor link, empty in 83 % of rows
source_filestringname of the original .html
extra_metastringJSON of any remaining front-matter fields (image, section, tags, …). Empty for most rows
textstringarticle body verbatim, all markdown markup preserved
n_charsint32length of text in characters
crlfboolthe source file used CRLF line endings (true for every file in this corpus)
utf8_okboolthe file decodes as strict UTF-8 (true for every file)
yaml_validboolthe front matter is valid YAML. false for 234 rows, where the fallback parser was used
sha1_bodystringSHA-1 of the normalized body — the key for grouping duplicates
sha1_rawstringSHA-1 of the original .md file as a whole. Verifies that packing lost nothing
dup_groupint320 = body is unique; otherwise the index of the canonical row (canonical = lexicographically smallest path)
id_repeatboolthe same message_id occurs more than once in the corpus (1,051 rows)

About text

The body is stored verbatim, with two normalizations, both reversible:

  1. 1.CRLF → LF (recorded in the crlf flag),
  2. 2.leading and trailing blank lines stripped (pure visual noise).

Line breaks inside paragraphs and all markup (#, **, >, tables, links, images) are exactly as in the original. The first lines almost always contain a duplicated # title and a *date · author* byline — they are already available as separate columns, and tools/load_example.py shows how to strip them before training.

Loading

python
from datasets import load_dataset

ds = load_dataset("rustemgareev/tg-contest", split="train")                   # full
ds = load_dataset("rustemgareev/tg-contest", split="train", streaming=True)   # streamed

# training text: body without the duplicated header
import re
def prep(row):
    text = re.sub(r"^#\s+.*\n+", "", row["text"], count=1)
    return re.sub(r"\n?\*[^*\n]{0,120}\*\n", "\n", text, count=1).strip()

# attach participant annotations
import pandas as pd
annot = pd.read_csv("annotations.csv")
labels = dict(zip(annot["id"], annot["label"]))
ds = ds.filter(lambda r: r["id"] in labels).map(lambda r: {**r, "label": labels[r["id"]]})

Files live at data/train-00000-of-00021-20200427.parquet … — they can also be read directly with pyarrow.parquet, without datasets.

Known limitations

  • —No labels. The corpus is unlabeled by design — participants annotate it. There is no label field in this dataset.
  • —`lang` is unreliable. It is a heuristic from the upstream converter, not ground truth. The converter also discarded everything it did not recognize as ru or en, so only those two languages remain and the size of the discarded portion is unknown. If language matters, re-detect it from text.
  • —The source channel is not in the corpus. The Telegram channel a post came from was never recorded: the filename is a bare message_id, and site is the article's outlet, not the channel. Grouping "by channel" is therefore impossible; grouping "by outlet" works.
  • —234 rows have invalid YAML in their front matter: the converter escaped only \ and ", but not line breaks. Values were recovered with a line-based fallback parser and are flagged yaml_valid=False.
  • —46 rows contain U+FFFD — a replacement character left over from conversion (utf8_ok=True; the replacement was already in the source).
  • —1,051 rows duplicate a `message_id` with different content; no two files are byte-identical. Check id_repeat when building splits.
  • —Duplicates were not removed. They amount to 0.43 %, so removal would gain almost nothing, while 5.4 % of normalized-title matches are mostly different articles with similar headlines and would have been discarded wrongly. Instead of deletion there are the dup_group and sha1_body columns — the decision is yours.

Integrity

bash
sha256sum -c stats/shards.sha256                 # Linux
Get-FileHash data\*.parquet -Algorithm SHA256   # Windows

stats/manifest.json holds per-day statistics: row counts, language distribution, broken front matter, missing fields, shard size. stats/duplicates.json and stats/duplicates.csv describe the duplicate groups.

Packing was verified by an end-to-end round-trip: 400 randomly sampled rows were re-read from the original .md files and compared on sha1_raw and every field — byte-identical, 400/400.

How this was built

Source: 790,166 files under YYYYMMDD/HH/<message_id>.md (21 days × 24 hours), 3.63 GB, CRLF, UTF-8, front matter present in every file.

  1. 1.build_parquet.py — a single pass: a pool of 16 threads doing read() only, parsing in the main process. The source sits on Storage Spaces backed by a single spinning HDD (~4 ms per file), so extra read threads do not help, while threads fighting over the GIL inside the regexes actively hurt. Writes one shard per day, then an atomic rename.
  2. 2.assign_dup_groups.py — labels duplicates using the finished shards, not a second pass over the HDD.
  3. 3.validate_parquet.py — schema, checksums, round-trip against the source files, and dup_group consistency.

Full pass over the corpus: 37 minutes; the whole pipeline including the second stage takes about 45 minutes. The scripts live in the project's git repository and are not published as part of the dataset.