Team Ai
Datasetpublic

jpaulpoliquit/ph-pretrain

PH Pretrain β€” Philippine Languages Corpus (v0.6-ph-unified) πŸ‘‰ Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead β€” it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes264downloads
Dataset Card

PH Pretrain β€” Philippine Languages Corpus (v0.6-ph-unified)

πŸ‘‰ Looking to train a Filipino/Tagalog model? Use [`jpaulpoliquit/ph-pretrain-03`](https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03) instead β€” it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.

A cleaned, deduplicated, document-level pretraining corpus for Philippine languages: Cebuano, Waray, Tagalog/Filipino, Ilocano, Bikol, Kapampangan, and Pangasinan.

  • β€”7,176,685 documents across train / validation / test
  • β€”β‰ˆ 1.7 billion tokens (Sailor2-1B tokenizer; see Size & tokens)
  • β€”~1.33 GB to download (Parquet, zstd) Β· ~8.2 GB uncompressed
  • β€”Built by the `jpaulpoliquit/pretraining` data refinery (12-stage filtering, dedup, and quality scoring)
Read this first. ~95.7% of documents are Cebuano and Waray Wikipedia, a large share of which is bot-generated (Lsjbot) and therefore short and templated. Treat language balance and register as something to re-weight for your use case β€” see Limitations & biases.

This is a single, dedup-safe merge of two refinery tracks (regional Wikipedia + a Filipino-forward web/curated mix) with 0 overlapping `document_id`s. Release manifest: `release_v0.6-ph-unified.json`.

Composition

By source (documents)

SourceDocumentsShare
wikimedia_ceb (Cebuano Wikipedia)5,932,31982.66%
wikimedia_war (Waray Wikipedia)935,37213.03%
fineweb2_tl_latn (Tagalog web, FineWeb2)205,9362.87%
wikimedia_tl (Tagalog Wikipedia)39,9510.56%
halohalo_combined (Filipino curated)31,9400.45%
wikimedia_ilo (Ilocano Wikipedia)12,0850.17%
wikimedia_bcl (Central Bikol Wikipedia)11,7410.16%
wikimedia_pam (Kapampangan Wikipedia)6,5410.09%
wikimedia_pag (Pangasinan Wikipedia)8000.01%

Cebuano + Waray Wikipedia = 95.69% of documents. Tagalog/Filipino (tl) across all sources = 277,827 docs (3.87%).

By register

RegisterDocuments
cebuano5,932,319
waray935,372
filipino_formal245,887
filipino_casual31,940
ilocano12,085
bicolano11,741
kapampangan6,541
pangasinan800

register is a coarse editorial label assigned by the refinery from source priors + heuristics, not a sociolinguistic ground truth.

Quality bands

Each document gets a heuristic quality_score (0–1, mean β‰ˆ 0.78) and an A/B/C/D quality_bucket:

BucketDocumentsShare
A3,269,54345.56%
B3,906,03854.43%
C1,0940.02%
D10<0.01%

Splits

SplitDocuments
train7,050,119
validation63,409
test63,157

Split labels are inherited per document from the upstream refinery mixes (β‰ˆ0.9% held out each for validation/test). They are random holdouts for convenience, not a curated, contamination-controlled benchmark.

Size & tokens

MetricValue
Documents7,176,685
Characters~4.92 B
Whitespace words~771 M
Tokens (Sailor2-1B tokenizer)β‰ˆ 1.69 B
Median document length~406 chars (~60–70 words)
Download size (Parquet)~1.33 GB
Uncompressed~8.2 GB

The token figure was measured by tokenizing an 8.8k-document sample with the `sail/Sailor2-1B` tokenizer (β‰ˆ 2.19 tokens/word, 0.345 tokens/char) and scaling by the exact corpus character count. For Cebuano/Waray a word β‰ˆ 2.2 subword tokens, so naive words Γ— 1.3 heuristics underestimate this corpus by ~40%.

Usage

python
from datasets import load_dataset

# Full dataset (β‰ˆ1.33 GB download)
ds = load_dataset("jpaulpoliquit/ph-pretrain")
row = ds["train"][0]
print(row["language_guess"], row["register"], row["text"][:300])

# Stream without downloading everything
stream = load_dataset("jpaulpoliquit/ph-pretrain", split="train", streaming=True)
for row in stream:
    text = row["text"]
    break

Filter to a language, register, or quality band:

python
tagalog = ds["train"].filter(lambda r: r["language_guess"] == "tl")
high_q  = ds["train"].filter(lambda r: r["quality_bucket"] == "A")

For pretraining you typically only need text (and maybe language_guess / quality_bucket for sampling weights).

Schema

Data are stored as flat Parquet at data/{train,validation,test}/part-*.parquet.

Core

ColumnTypeNotes
document_idstringStable 20-char id; unique across the dataset
textstringFull document body (UTF-8)
text_sha256string64-char content hash (exact-dedup key)

Provenance

ColumnTypeNotes
source_namestringe.g. wikimedia_ceb, fineweb2_tl_latn
source_urlstringArticle / source URL when available
source_typestringwikimedia_dump or hf_curated
source_reliabilityfloatSource prior (0.95 for Wikimedia)
crawl_datestringISO timestamp of refinery ingest
licensestringAuthoritative per-document license (cc-by-sa-4.0 or odc-by-1.0)
redistributeboolRedistribution allowed (all true here)
train_onlyboolTrain-only restriction (all false here)

Language & register

ColumnTypeNotes
language_guessstringISO 639-3 code from language ID
language_bucketstringCoarse language bucket
registerstringEditorial register label (see above)

Quality

ColumnTypeNotes
quality_scorefloat0–1 heuristic score
quality_bucketstringA / B / C / D
is_syntheticboolSynthetic-text flag (all false here)

`metadata_json` (string) carries the remaining refinery fields as a JSON blob.

Field population notes (this release)

  • β€”language_confidence is empty (null) for every row in this release β€” use quality_score and language_guess instead.
  • β€”dialect, created_date, contamination_verdict, and code_mix_type are empty in this release (the pipeline stages that fill them did not run on this mix).
  • β€”is_synthetic is false and train_only is false for all rows.

How it was built

Documents flow through the refinery's typed pipeline (full source: `jpaulpoliquit/pretraining`):

  1. 1.Ingest raw Wikimedia dumps and Hugging Face sources.
  2. 2.Language ID + register assignment.
  3. 3.Quality filtering & scoring β€” short, boilerplate, and low-signal documents are dropped (text floor β‰ˆ 200 chars) and the rest are scored.
  4. 4.Exact dedup by text_sha256.
  5. 5.Mix & split into train/validation/test.
  6. 6.Unify the regional and Filipino tracks, deduplicating by document_id (0 cross-track collisions).

Licensing

This corpus mixes two licenses; the per-document license column is authoritative.

LicenseDocumentsShareSources
CC-BY-SA-4.06,938,80996.69%All Wikimedia dumps (wikimedia_*)
ODC-BY-1.0237,8763.31%fineweb2_tl_latn, halohalo_combined

The dataset is tagged CC-BY-SA-4.0 because the copyleft Wikimedia portion dominates and is the most restrictive term. If you redistribute the Wikimedia-derived rows you must comply with CC-BY-SA-4.0 (attribution + share-alike); the ODC-BY rows require attribution. Underlying web text (FineWeb2) derives from Common Crawl and remains subject to original publishers' rights. You are responsible for your own license compliance.

Intended use

  • β€”Continued pretraining / domain adaptation of language models for Philippine languages, especially the low-resource regional ones.
  • β€”Tokenizer training and language-coverage analysis for the Philippines.
  • β€”A base to re-weight or subset (e.g. downsample Cebuano/Waray Wikipedia, upweight Tagalog web) for a more balanced mix.

Limitations & biases

  • β€”Heavy Cebuano/Waray Wikipedia skew + bot content. ~95.7% of documents are ceb/war Wikipedia, much of it created by the Lsjbot bot from structured templates. These articles are short, formulaic, and highly repetitive. Training directly on the raw mix will over-represent this style; consider downsampling or quality/length filtering.
  • β€”Tagalog/Filipino is a small slice (~3.9%) despite being the most widely spoken language β€” supplement if Tagalog is your target.
  • β€”Register and language labels are heuristic, not gold. Expect some misclassification, especially among closely related languages.
  • β€”No contamination filtering recorded in this release (contamination_verdict is empty); the train/val/test split is a random holdout, not a leakage-controlled benchmark.
  • β€”Smallest languages are tiny (Pangasinan = 800 docs) and not sufficient on their own.

Provenance & reproducibility

  • β€”Manifest (source of truth): `release_v0.6-ph-unified.json` in this repo.
  • β€”Pipeline & build scripts: github.com/jpaulpoliquit/pretraining
  • β€”Release version: v0.6-ph-unified

Citation

bibtex
@misc{ph_pretrain_v06_unified,
  title        = {PH Pretrain: A Philippine Languages Pretraining Corpus (v0.6-ph-unified)},
  author       = {Poliquit, John Paul},
  year         = {2026},
  howpublished = {Hugging Face Datasets},
  url          = {https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain}
}

Please also cite the upstream sources: Wikimedia (CC-BY-SA-4.0) and FineWeb2 (ODC-BY-1.0).