oliveirabruno01/ptbr-creative-cpt
PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.
PT-BR Creative Corpus v0.1.0
A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research.
Status
This is the canonical corpus freeze, not a final model-specific training build.
- Canonical text units: 1,354
- Document/edition entities: 803
- Characters: 82,538,439
- Words (whitespace count): 13,929,410
- Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the
audit_metricsconfig.
The canonical corpus config deliberately contains no generic token-count column. Exact token counts belong to tokenizer-specific CPT builds.
Load a small sample without loading the corpus
from datasets import load_dataset
ds = load_dataset("oliveirabruno01/ptbr-creative-cpt", "corpus", split="corpus", streaming=True)
for row in ds.take(100):
...Load the full canonical corpus
from datasets import load_dataset
ds = load_dataset("oliveirabruno01/ptbr-creative-cpt", "corpus", split="corpus")For CPT, derive a model/run-specific training build from this corpus:
canonical corpus
-> work/author-level benchmark holdout + decontamination
-> target model tokenizer
-> EOS/document-boundary policy
-> sequence packing for chosen context length
-> trainer-specific binary/memmap formatTokenized/packed sequences are intentionally not baked into this repository.
Text-unit model
Most corpus rows are complete selected works/documents. FE-Unicamp is intentionally different: modern scholarly editions were segmented and only blocks classified as primary historical creative literature were retained. Those rows use unit_kind="selected_segment" and point to a real parent edition through parent_document_id.
For FE, primary_authors refers to the historical literary writer; modern scholarly editors are kept separately in editors.
Orthography
text_variant="original" means the extracted source orthography is preserved. Any future AO90-normalized material must be released as a separate derived variant, never as a replacement for original text.
Rights
Rights vary by source and are recorded per row with rights_tier, license, license_url, and rights_evidence.
Important caveat: the ptbr-books-publicos backbone is upstream-released under CC0-1.0 and described as public-domain, but this project did not independently establish per-item rights for every backbone book. The corpus therefore records that as a source-dataset claim rather than silently converting it into per-item legal provenance.
BBM scan-derived text is not included in corpus; it remains in the review/deferred inventory.
Auxiliary configs
documents: document/edition hierarchy.review_inventory: deferred and review queues, without training text.curation_lineage: acceptance/dedup/curation lineage for canonical units.audit_metrics: historical project accounting metrics. These are audit artifacts, not canonical token counts.
Data architecture references
This repository keeps canonical text separate from tokenizer-specific training preprocessing, following the general architecture used by open pretraining projects such as Dolma/OLMo.
- Dolma data format: https://github.com/allenai/dolma/blob/main/docs/data-format.md
- OLMo memmap preprocessing: https://github.com/allenai/OLMo/blob/main/scripts/preparememmapdataset.py
- Hugging Face manual dataset configs: https://huggingface.co/docs/hub/datasets-manual-configuration
