exoarbuus/goatis-transcripts
Goatis / Sv3rige Video Transcripts Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net. What makes it more than raw ASR Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.
Goatis / Sv3rige Video Transcripts
Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net.
What makes it more than raw ASR
Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a naive transcript attributes other people's words to him. Here, every diarized speaker cluster carries a goatis_score — cosine similarity of its ECAPA voice embedding against reference centroids built from five known-solo videos spanning 2012–2026. Words from reacted-to videos are separable from his own speech.
Schema (videos.jsonl.gz, one row per video)
Pipeline
WhisperX (Whisper large-v3, float16, English) → wav2vec2 forced alignment (word timestamps) → pyannote 3.1 diarization → speechbrain ECAPA embeddings vs reference centroids. Processed on 2× RTX 5090 in ~18 GPU-hours. Total compute cost of the corpus: about $3.
Usage
from datasets import load_dataset
ds = load_dataset("json", data_files="hf://datasets/exoarbuus/goatis-transcripts/videos.jsonl.gz")
# his own words only, per video
for row in ds["train"]:
his = {s for s, v in row["speakers"].items() if v["goatis_score"] >= 0.55}
text = " ".join(seg["text"] for seg in row["segments"] if seg["speaker"] in his)Caveats
- ASR output: proper nouns are the weak point (e.g. "Aajonus" is often mangled in speech); numbers and common speech are reliable.
goatis_scoreis per cluster, not per word; overlapping speech and brief interjections can be misattributed. Scores near the 0.55 threshold are genuinely ambiguous — treat them as such.- ~46 videos are music/compilations where no speaker matches him (correctly).
- Sarcasm, quoted speech, and read-aloud comments are his voice and are attributed to him: vocal attribution ≠ endorsement of the words.
Provenance & license
Transcripts of publicly posted YouTube videos, provided for archival and research use. Underlying speech remains the speaker's. No warranty of transcription accuracy; verify against the linked timestamps before quoting.
