jpaulpoliquit/ph-pretrain
PH Pretrain β Philippine Languages Corpus (v0.6-ph-unified) π Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead β it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus forβ¦ See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.
PH Pretrain β Philippine Languages Corpus (v0.6-ph-unified)
π Looking to train a Filipino/Tagalog model? Use [`jpaulpoliquit/ph-pretrain-03`](https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03) instead β it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.A cleaned, deduplicated, document-level pretraining corpus for Philippine languages: Cebuano, Waray, Tagalog/Filipino, Ilocano, Bikol, Kapampangan, and Pangasinan.
- 7,176,685 documents across train / validation / test
- β 1.7 billion tokens (Sailor2-1B tokenizer; see Size & tokens)
- ~1.33 GB to download (Parquet, zstd) Β· ~8.2 GB uncompressed
- Built by the `jpaulpoliquit/pretraining` data refinery (12-stage filtering, dedup, and quality scoring)
Read this first. ~95.7% of documents are Cebuano and Waray Wikipedia, a large share of which is bot-generated (Lsjbot) and therefore short and templated. Treat language balance and register as something to re-weight for your use case β see Limitations & biases.
This is a single, dedup-safe merge of two refinery tracks (regional Wikipedia + a Filipino-forward web/curated mix) with 0 overlapping `document_id`s. Release manifest: `release_v0.6-ph-unified.json`.
Composition
By source (documents)
Cebuano + Waray Wikipedia = 95.69% of documents. Tagalog/Filipino (tl) across all sources = 277,827 docs (3.87%).
By register
register is a coarse editorial label assigned by the refinery from source priors + heuristics, not a sociolinguistic ground truth.
Quality bands
Each document gets a heuristic quality_score (0β1, mean β 0.78) and an A/B/C/D quality_bucket:
Splits
Split labels are inherited per document from the upstream refinery mixes (β0.9% held out each for validation/test). They are random holdouts for convenience, not a curated, contamination-controlled benchmark.
Size & tokens
The token figure was measured by tokenizing an 8.8k-document sample with the `sail/Sailor2-1B` tokenizer (β 2.19 tokens/word, 0.345 tokens/char) and scaling by the exact corpus character count. For Cebuano/Waray a word β 2.2 subword tokens, so naive words Γ 1.3 heuristics underestimate this corpus by ~40%.
Usage
from datasets import load_dataset
# Full dataset (β1.33 GB download)
ds = load_dataset("jpaulpoliquit/ph-pretrain")
row = ds["train"][0]
print(row["language_guess"], row["register"], row["text"][:300])
# Stream without downloading everything
stream = load_dataset("jpaulpoliquit/ph-pretrain", split="train", streaming=True)
for row in stream:
text = row["text"]
breakFilter to a language, register, or quality band:
tagalog = ds["train"].filter(lambda r: r["language_guess"] == "tl")
high_q = ds["train"].filter(lambda r: r["quality_bucket"] == "A")For pretraining you typically only need text (and maybe language_guess / quality_bucket for sampling weights).
Schema
Data are stored as flat Parquet at data/{train,validation,test}/part-*.parquet.
Core
Provenance
Language & register
Quality
`metadata_json` (string) carries the remaining refinery fields as a JSON blob.
Field population notes (this release)
language_confidenceis empty (null) for every row in this release β usequality_scoreandlanguage_guessinstead.dialect,created_date,contamination_verdict, andcode_mix_typeare empty in this release (the pipeline stages that fill them did not run on this mix).is_syntheticisfalseandtrain_onlyisfalsefor all rows.
How it was built
Documents flow through the refinery's typed pipeline (full source: `jpaulpoliquit/pretraining`):
- Ingest raw Wikimedia dumps and Hugging Face sources.
- Language ID + register assignment.
- Quality filtering & scoring β short, boilerplate, and low-signal documents are dropped (text floor β 200 chars) and the rest are scored.
- Exact dedup by
text_sha256. - Mix & split into train/validation/test.
- Unify the regional and Filipino tracks, deduplicating by
document_id(0 cross-track collisions).
Licensing
This corpus mixes two licenses; the per-document license column is authoritative.
The dataset is tagged CC-BY-SA-4.0 because the copyleft Wikimedia portion dominates and is the most restrictive term. If you redistribute the Wikimedia-derived rows you must comply with CC-BY-SA-4.0 (attribution + share-alike); the ODC-BY rows require attribution. Underlying web text (FineWeb2) derives from Common Crawl and remains subject to original publishers' rights. You are responsible for your own license compliance.
Intended use
- Continued pretraining / domain adaptation of language models for Philippine languages, especially the low-resource regional ones.
- Tokenizer training and language-coverage analysis for the Philippines.
- A base to re-weight or subset (e.g. downsample Cebuano/Waray Wikipedia, upweight Tagalog web) for a more balanced mix.
Limitations & biases
- Heavy Cebuano/Waray Wikipedia skew + bot content. ~95.7% of documents are
ceb/warWikipedia, much of it created by the Lsjbot bot from structured templates. These articles are short, formulaic, and highly repetitive. Training directly on the raw mix will over-represent this style; consider downsampling or quality/length filtering. - Tagalog/Filipino is a small slice (~3.9%) despite being the most widely spoken language β supplement if Tagalog is your target.
- Register and language labels are heuristic, not gold. Expect some misclassification, especially among closely related languages.
- No contamination filtering recorded in this release (
contamination_verdictis empty); the train/val/test split is a random holdout, not a leakage-controlled benchmark. - Smallest languages are tiny (Pangasinan = 800 docs) and not sufficient on their own.
Provenance & reproducibility
- Manifest (source of truth): `release_v0.6-ph-unified.json` in this repo.
- Pipeline & build scripts: github.com/jpaulpoliquit/pretraining
- Release version:
v0.6-ph-unified
Citation
@misc{ph_pretrain_v06_unified,
title = {PH Pretrain: A Philippine Languages Pretraining Corpus (v0.6-ph-unified)},
author = {Poliquit, John Paul},
year = {2026},
howpublished = {Hugging Face Datasets},
url = {https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain}
}Please also cite the upstream sources: Wikimedia (CC-BY-SA-4.0) and FineWeb2 (ODC-BY-1.0).
