Team Ai
Datasetpublic

oliveirabruno01/ptbr-creative-cpt

PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.

sourceHugging Faceotherupdated 19d agoView on Hugging Face
0likes77downloads
Dataset Card

PT-BR Creative Corpus v0.1.0

A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research.

Status

This is the canonical corpus freeze, not a final model-specific training build.

  • —Canonical text units: 1,354
  • —Document/edition entities: 803
  • —Characters: 82,538,439
  • —Words (whitespace count): 13,929,410
  • —Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config.

The canonical corpus config deliberately contains no generic token-count column. Exact token counts belong to tokenizer-specific CPT builds.

Load a small sample without loading the corpus

python
from datasets import load_dataset

ds = load_dataset("oliveirabruno01/ptbr-creative-cpt", "corpus", split="corpus", streaming=True)
for row in ds.take(100):
    ...

Load the full canonical corpus

python
from datasets import load_dataset

ds = load_dataset("oliveirabruno01/ptbr-creative-cpt", "corpus", split="corpus")

For CPT, derive a model/run-specific training build from this corpus:

text
canonical corpus
  -> work/author-level benchmark holdout + decontamination
  -> target model tokenizer
  -> EOS/document-boundary policy
  -> sequence packing for chosen context length
  -> trainer-specific binary/memmap format

Tokenized/packed sequences are intentionally not baked into this repository.

Text-unit model

Most corpus rows are complete selected works/documents. FE-Unicamp is intentionally different: modern scholarly editions were segmented and only blocks classified as primary historical creative literature were retained. Those rows use unit_kind="selected_segment" and point to a real parent edition through parent_document_id.

For FE, primary_authors refers to the historical literary writer; modern scholarly editors are kept separately in editors.

Orthography

text_variant="original" means the extracted source orthography is preserved. Any future AO90-normalized material must be released as a separate derived variant, never as a replacement for original text.

Rights

Rights vary by source and are recorded per row with rights_tier, license, license_url, and rights_evidence.

Important caveat: the ptbr-books-publicos backbone is upstream-released under CC0-1.0 and described as public-domain, but this project did not independently establish per-item rights for every backbone book. The corpus therefore records that as a source-dataset claim rather than silently converting it into per-item legal provenance.

BBM scan-derived text is not included in corpus; it remains in the review/deferred inventory.

Auxiliary configs

  • —documents: document/edition hierarchy.
  • —review_inventory: deferred and review queues, without training text.
  • —curation_lineage: acceptance/dedup/curation lineage for canonical units.
  • —audit_metrics: historical project accounting metrics. These are audit artifacts, not canonical token counts.

Data architecture references

This repository keeps canonical text separate from tokenizer-specific training preprocessing, following the general architecture used by open pretraining projects such as Dolma/OLMo.

  • —Dolma data format: https://github.com/allenai/dolma/blob/main/docs/data-format.md
  • —OLMo memmap preprocessing: https://github.com/allenai/OLMo/blob/main/scripts/preparememmapdataset.py
  • —Hugging Face manual dataset configs: https://huggingface.co/docs/hub/datasets-manual-configuration