Team Ai
Datasetpublic

paulalesius/terence-mckenna-transcripts

Terence McKenna Transcripts Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables: talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags) turns — one row per non-empty line; speaker tags stripped from text Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column. Load from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes20downloads
Dataset Card

Terence McKenna Transcripts

Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables:

  • —`talks` — one row per video; full original transcript (including [SPEAKER_XX] diarization tags)
  • —`turns` — one row per non-empty line; speaker tags stripped from text

Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column.

Load

python
from datasets import load_dataset

talks = load_dataset("paulalesius/terence-mckenna-transcripts", "talks")
turns = load_dataset("paulalesius/terence-mckenna-transcripts", "turns")

Configs

talks

columntypedescription
idstringVideo title slug (filename without .transcript.txt)
titlestringSame as id, lightly cleaned
filenamestringOriginal source filename
transcriptstringFull file text, including speaker tags and line breaks
num_charsintlen(transcript)
num_turnsintNumber of turns produced from this file

172 rows.

turns

columntypedescription
idstringSame id as the parent talk
titlestringSame title as the parent talk
turn_indexint0-based line index within that talk
textstringOne transcript line, [SPEAKER_XX] prefix removed

135,672 rows.

A line like:

text
[SPEAKER_03] Why are we willing to go along with this shell game?

is stored in turns as text = "Why are we willing to go along with this shell game?" and remains tagged in talks.transcript.

Size (from the current build)

Talks172
Turns135,672
Characters in talks.transcript14,068,290
Chars / talkmin 821, median ~60k, max ~476k
Turns / talkmin 10, median ~549, max 5,000

Long workshops dominate the upper end (e.g. Search For The Original Tree Of Knowledge, Esalen residencies). A handful of files are very short clips or fragments.

Source and processing

  • —Source: 172 video recordings; each transcript file was named after the video.
  • —ASR + local speaker diarization produced lines of the form [SPEAKER_NN] utterance.
  • —SPEAKER_NN is per file only. The same label in two files is not the same person.
  • —Empty lines and tag-only lines were dropped from turns.
  • —No cross-video deduplication was applied. Some titles look like the same talk uploaded twice (e.g. overlapping “Tryptamine Hallucinogens”, “Definitive UFO Tape”, “New & Old Maps of Hyperspace”, “Rap Dancing” recordings).
  • —The catalog includes some Dennis McKenna talks and multi-speaker trialogues (McKenna / Sheldrake / Abraham), not only Terence solo lectures.

ASR errors are present (misheard names, broken sentences, missing punctuation).

Intended use

Research, search, citation, language modeling, and retrieval over this lecture archive.

Not intended as a verified scholarly edition of McKenna’s words.

Limitations

  • —Automatic speech recognition, not human-corrected transcripts.
  • —No timestamps.
  • —No global speaker identities.
  • —Possible duplicate recordings.
  • —Mix of Terence McKenna, Dennis McKenna, interviewers, and panelists.
  • —Very short files may be trailers, readings, or incomplete jobs.

License and copyright

This repo packages unofficial machine transcripts I made from publicly circulated videos.

  • —I am not the copyright holder of the talks, interviews, or readings.
  • —I am not affiliated with the McKenna estate, publishers, or original venues.
  • —I do not grant a license to Terence McKenna’s (or Dennis McKenna’s, or any guest’s) words.
  • —What I did — diarized ASR, filenames aligned to the videos, and this table layout — you may reuse for research and model work.
  • —[SPEAKER_XX] labels are local to each file. They are not voice IDs and do not identify anyone across talks.

If a rights holder objects, the files should come down. If you use the corpus, you are responsible for how you use it. This is a research convenience, not an official text.

Citation

bibtex
@misc{mckenna_transcripts,
  title  = {Terence McKenna Transcripts},
  author = {Paul Alesius},
  year   = {2026},
  url    = {https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts}
}