Team Ai
Datasetpublic

exoarbuus/goatis-transcripts

Goatis / Sv3rige Video Transcripts Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net. What makes it more than raw ASR Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes12downloads
Dataset Card

Goatis / Sv3rige Video Transcripts

Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net.

What makes it more than raw ASR

Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a naive transcript attributes other people's words to him. Here, every diarized speaker cluster carries a goatis_score — cosine similarity of its ECAPA voice embedding against reference centroids built from five known-solo videos spanning 2012–2026. Words from reacted-to videos are separable from his own speech.

Schema (videos.jsonl.gz, one row per video)

fielddescription
video_idYouTube id — https://youtube.com/watch?v=<id>
title, channel, upload_date, url, view_countvideo metadata
duration_secaudio duration
speakersper diarized speaker: talk_sec, goatis_score (cosine, ~0.55+ = him)
segments[]start, end, text, speaker, words[]
segments[].words[][word, start, end, speaker, align_score] — word-level timestamps

Pipeline

WhisperX (Whisper large-v3, float16, English) → wav2vec2 forced alignment (word timestamps) → pyannote 3.1 diarization → speechbrain ECAPA embeddings vs reference centroids. Processed on 2× RTX 5090 in ~18 GPU-hours. Total compute cost of the corpus: about $3.

Usage

python
from datasets import load_dataset
ds = load_dataset("json", data_files="hf://datasets/exoarbuus/goatis-transcripts/videos.jsonl.gz")

# his own words only, per video
for row in ds["train"]:
    his = {s for s, v in row["speakers"].items() if v["goatis_score"] >= 0.55}
    text = " ".join(seg["text"] for seg in row["segments"] if seg["speaker"] in his)

Caveats

  • —ASR output: proper nouns are the weak point (e.g. "Aajonus" is often mangled in speech); numbers and common speech are reliable.
  • —goatis_score is per cluster, not per word; overlapping speech and brief interjections can be misattributed. Scores near the 0.55 threshold are genuinely ambiguous — treat them as such.
  • —~46 videos are music/compilations where no speaker matches him (correctly).
  • —Sarcasm, quoted speech, and read-aloud comments are his voice and are attributed to him: vocal attribution ≠ endorsement of the words.

Provenance & license

Transcripts of publicly posted YouTube videos, provided for archival and research use. Underlying speech remains the speaker's. No warranty of transcription accuracy; verify against the linked timestamps before quoting.