Team Ai
Datasetpublic

nlp-chula/ThaiTrees

ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.

sourceHugging Facecc-by-sa-4.0updated 17d agoView on Hugging Face
2likes610downloads
Dataset Card

ThaiTrees

A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U.

Dataset Summary

ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared across all three artefacts, so results can be joined at the document level. Token identifiers do not align across the lexicon and the parsed corpus, which index different tokenisations of the same raw text (see Processing Pipeline).

  • —Raw text corpus — 366,120 documents and 341,967,133 AttaCut-segmented tokens, released as Parquet.
  • —Frequency lexicon — 451,252 unique word forms and 255,434,917 filtered tokens, with per-domain count, rank, frequency per million, document frequency and IDF. Covers the full raw corpus.
  • —Dependency-parsed corpus — 199,836,464 dependency edges over 203,892,200 tokens in 10-column CoNLL-U, with PhayaThaiBERT POS tags.

Thai already has a manually annotated dependency treebank, UD Thai-TUD (3,627 dependency trees, 77,215 tokens). ThaiTrees complements rather than replaces it: ThaiTrees is automatically parsed and not validated against human annotation, so use UD Thai-TUD as the gold standard for parser evaluation.

Languages

Thai (th).

Dataset Structure

Configurations

ConfigDescriptionSplits
rawRaw text corpusnews, wikipedia, spoken, social_media
lexiconFrequency lexicontrain (single split)
parsedDependency-parsed CoNLL-Unews, wikipedia, spoken, social_media

Data Fields

`raw` configuration (one row per document):

FieldTypeDescription
domainstringOne of news / wikipedia / spoken / social_media
idstringDocument identifier (e.g. news_1); matches doc_id in parsed
urlstringSource URL
titlestringDocument title
categorieslist[string]Source categories or tags, where available (may be null)
textstringThe original, unsegmented document text
segmented_textstringThe document text, AttaCut-tokenised, with `\` between tokens
token_countint64Number of AttaCut tokens in the document

`lexicon` configuration (one row per unique word form, 34 columns):

FieldTypeDescription
wordstringThe word form
total_countintCross-corpus occurrence count
total_rankintCross-corpus rank by frequency
total_freq_per_millionfloatFrequency per million tokens, cross-corpus
total_doc_countintNumber of documents (corpus-wide) containing this word
{domain}_countintOccurrence count in {domain}
{domain}_rankintWithin-domain rank
{domain}_freq_per_millionfloatFrequency per million tokens within domain
{domain}_doc_countintNumber of documents in {domain} containing this word
{domain}_doc_freqfloatDocument frequency in {domain} (0–1)
{domain}_idffloatIDF computed over {domain} documents
domain_countintHow many of the four domains the word appears in
is_cross_domainboolTrue if the word appears in all four domains
word_lengthintCharacter length of the word
idffloatCross-corpus IDF
max_domainstringThe domain in which the word's freq/M is highest

A SQLite copy of the lexicon (lexicon/lexicon.sqlite) is also provided.

`parsed` configuration (one row per sentence):

FieldTypeDescription
doc_idstringDocument identifier (matches id in raw)
domainstringDomain tag
sent_iduint32Sentence index within the document
tokenslist[string]FORM (CoNLL-U column 2)
lemmalist[string]LEMMA (column 3); placeholder, always .
uposlist[string]UPOS (column 4); PhayaThaiBERT POS tags
xposlist[string]XPOS (column 5); placeholder, always .
featslist[string]FEATS (column 6); placeholder, always _
headlist[int32]HEAD (column 7); 1-indexed, 0 = root
deprellist[string]DEPREL (column 8); Universal Dependencies relation
depslist[string]DEPS (column 9); placeholder, always _
misclist[string]MISC (column 10); character offsets

Data Splits

SplitDocumentsTokensSentences
news100,84652,641,127897,348
wikipedia175,069101,378,5831,303,443
spoken2,50530,238,809266,944
social_media87,700157,708,6142,052,040
Total366,120341,967,1334,519,775

The parsed corpus contains 203,892,200 tokens and 199,836,464 dependency edges. The lexicon config has a single train split with 451,252 rows (one per unique word form; see Processing Pipeline).

Dataset Creation

Curation Rationale

Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank (UD Thai-TUD) for training and evaluating parsers, but lacked a large automatically parsed corpus for quantitative syntactic research. ThaiTrees fills that gap: a four-domain corpus whose sentences are parsed automatically under Universal Dependencies, released with a frequency lexicon in machine-readable formats.

Source Data

DomainSourceCollection
newsThaiPBS public-broadcaster websiteArticles publicly available on the website in December 2025
wikipediaThai WikipediaMediaWiki random-article endpoint, namespace 0, batches of 500, disambiguation pages excluded, reference sections truncated; sampled December 2025
spokenYouTube: Thai PBS Podcast network, independent podcasts (bigboung, BeSider), film subtitles, Prachatai political-commentary livestreams, SaltymanTH livestreamsyoutube-transcript.io API; channels selected manually for their manual transcription
social_mediaWisesight sentiment dataset and Pantip forumThe same sources used in training WangchanBERTa and PhayaThaiBERT

Processing Pipeline

The lexicon and the parsed corpus are produced in separate branches that apply word segmentation to different units of text: whole documents for the lexicon, individual sentences for parsing.

Frequency lexicon. AttaCut (Chormai et al., COLING 2020), a CNN-based Thai word segmenter, segments each whole document. Lexicon construction then removes punctuation, pure-numeric tokens, emoji, whitespace-only tokens, and forms whose total corpus frequency is at most five. After filtering, 255,434,917 tokens remain across 451,252 unique forms; 39,465 forms (8.7%) appear in all four domains.

Dependency-parsed corpus.

  1. 1.Sentence preparation. The raw text has newlines and runs of spaces collapsed to a single space, and social-media documents are capped at 500,000 characters (this truncates 149 long, concatenated documents).
  2. 2.Sentence segmentation with PyThaiNLP CRFcut, followed by AttaCut re-applied to each sentence string.
  3. 3.Dependency parsing with AttaParse 1.0, a Thai-specific wrapper around Stanza, on the pre-tokenised input. Whitespace-only tokens are discarded at the parser's pre-tokenised entry point. One deterministic parser failure on Wikipedia chunk 16 is excluded from the release.
  4. 4.POS tagging with a PhayaThaiBERT model fine-tuned on UD Thai-TUD (nlp-chula/phayathaibert-thai-pos-tagger), which achieves 90.64% accuracy and 81.34% macro F1 on the UD Thai-TUD test set. Its predictions overwrite the UPOS column (column 4) of the CoNLL-U output.

Additional Information

Dataset Curators

Attapol T. Rutherford and Papatchol Thientong, Department of Linguistics, Chulalongkorn University. Code: <https://github.com/nlp-chula/thaitrees>.

Citation Information

bibtex
@misc{rutherford2026thaitreesthaisyntacticdependency,
      title={ThaiTrees: Thai Syntactic Dependency Trees Across Domains}, 
      author={Attapol T. Rutherford and Papatchol Thientong},
      year={2026},
      eprint={2609.27558},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.27558}, 
}

Loading Examples

python
from datasets import load_dataset

# Raw text, per-domain split
ds_news = load_dataset("nlp-chula/ThaiTrees", "raw", split="news")
print(ds_news[0])
# {'domain': 'news', 'id': 'news_1', 'url': 'https://www.thaipbs.or.th/news/content/1',
#  'title': 'นักวิจัยออสเตรเลียเผยสาเหตุฉลามโจมตีมนุษย์', ...,
#  'segmented_text': 'นัก|วิจัย|ออสเตรเลีย|เผย|...', 'token_count': 230}

# Frequency lexicon (single split)
ds_lex = load_dataset("nlp-chula/ThaiTrees", "lexicon", split="train")
print(ds_lex[0])
# {'word': 'ที่', 'total_count': 6300593, 'total_rank': 1, ...}

# Dependency-parsed CoNLL-U, per-domain split
ds_parsed = load_dataset("nlp-chula/ThaiTrees", "parsed", split="news")
print(ds_parsed[0])
# {'doc_id': 'news_5028', 'domain': 'news', 'sent_id': 0,
#  'tokens': ['ย้าย', 'ผกก.', 'มีนบุรี', '-', 'หนองจอก', ...],
#  'upos': ['VERB', 'NOUN', 'PROPN', 'PUNCT', 'PROPN', ...], ...}