nlp-chula/ThaiTrees
ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.
ThaiTrees
A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U.
Dataset Summary
ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared across all three artefacts, so results can be joined at the document level. Token identifiers do not align across the lexicon and the parsed corpus, which index different tokenisations of the same raw text (see Processing Pipeline).
- Raw text corpus — 366,120 documents and 341,967,133 AttaCut-segmented tokens, released as Parquet.
- Frequency lexicon — 451,252 unique word forms and 255,434,917 filtered tokens, with per-domain count, rank, frequency per million, document frequency and IDF. Covers the full raw corpus.
- Dependency-parsed corpus — 199,836,464 dependency edges over 203,892,200 tokens in 10-column CoNLL-U, with PhayaThaiBERT POS tags.
Thai already has a manually annotated dependency treebank, UD Thai-TUD (3,627 dependency trees, 77,215 tokens). ThaiTrees complements rather than replaces it: ThaiTrees is automatically parsed and not validated against human annotation, so use UD Thai-TUD as the gold standard for parser evaluation.
Languages
Thai (th).
Dataset Structure
Configurations
Data Fields
`raw` configuration (one row per document):
`lexicon` configuration (one row per unique word form, 34 columns):
A SQLite copy of the lexicon (lexicon/lexicon.sqlite) is also provided.
`parsed` configuration (one row per sentence):
Data Splits
The parsed corpus contains 203,892,200 tokens and 199,836,464 dependency edges. The lexicon config has a single train split with 451,252 rows (one per unique word form; see Processing Pipeline).
Dataset Creation
Curation Rationale
Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank (UD Thai-TUD) for training and evaluating parsers, but lacked a large automatically parsed corpus for quantitative syntactic research. ThaiTrees fills that gap: a four-domain corpus whose sentences are parsed automatically under Universal Dependencies, released with a frequency lexicon in machine-readable formats.
Source Data
Processing Pipeline
The lexicon and the parsed corpus are produced in separate branches that apply word segmentation to different units of text: whole documents for the lexicon, individual sentences for parsing.
Frequency lexicon. AttaCut (Chormai et al., COLING 2020), a CNN-based Thai word segmenter, segments each whole document. Lexicon construction then removes punctuation, pure-numeric tokens, emoji, whitespace-only tokens, and forms whose total corpus frequency is at most five. After filtering, 255,434,917 tokens remain across 451,252 unique forms; 39,465 forms (8.7%) appear in all four domains.
Dependency-parsed corpus.
- Sentence preparation. The raw text has newlines and runs of spaces collapsed to a single space, and social-media documents are capped at 500,000 characters (this truncates 149 long, concatenated documents).
- Sentence segmentation with PyThaiNLP CRFcut, followed by AttaCut re-applied to each sentence string.
- Dependency parsing with AttaParse 1.0, a Thai-specific wrapper around Stanza, on the pre-tokenised input. Whitespace-only tokens are discarded at the parser's pre-tokenised entry point. One deterministic parser failure on Wikipedia chunk 16 is excluded from the release.
- POS tagging with a PhayaThaiBERT model fine-tuned on UD Thai-TUD (nlp-chula/phayathaibert-thai-pos-tagger), which achieves 90.64% accuracy and 81.34% macro F1 on the UD Thai-TUD test set. Its predictions overwrite the UPOS column (column 4) of the CoNLL-U output.
Additional Information
Dataset Curators
Attapol T. Rutherford and Papatchol Thientong, Department of Linguistics, Chulalongkorn University. Code: <https://github.com/nlp-chula/thaitrees>.
Citation Information
@misc{rutherford2026thaitreesthaisyntacticdependency,
title={ThaiTrees: Thai Syntactic Dependency Trees Across Domains},
author={Attapol T. Rutherford and Papatchol Thientong},
year={2026},
eprint={2609.27558},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.27558},
}Loading Examples
from datasets import load_dataset
# Raw text, per-domain split
ds_news = load_dataset("nlp-chula/ThaiTrees", "raw", split="news")
print(ds_news[0])
# {'domain': 'news', 'id': 'news_1', 'url': 'https://www.thaipbs.or.th/news/content/1',
# 'title': 'นักวิจัยออสเตรเลียเผยสาเหตุฉลามโจมตีมนุษย์', ...,
# 'segmented_text': 'นัก|วิจัย|ออสเตรเลีย|เผย|...', 'token_count': 230}
# Frequency lexicon (single split)
ds_lex = load_dataset("nlp-chula/ThaiTrees", "lexicon", split="train")
print(ds_lex[0])
# {'word': 'ที่', 'total_count': 6300593, 'total_rank': 1, ...}
# Dependency-parsed CoNLL-U, per-domain split
ds_parsed = load_dataset("nlp-chula/ThaiTrees", "parsed", split="news")
print(ds_parsed[0])
# {'doc_id': 'news_5028', 'domain': 'news', 'sent_id': 0,
# 'tokens': ['ย้าย', 'ผกก.', 'มีนบุรี', '-', 'หนองจอก', ...],
# 'upos': ['VERB', 'NOUN', 'PROPN', 'PUNCT', 'PROPN', ...], ...}