nht10/cx_sampled
CulturaX sampled text pools Use UPLOAD_COMPLETE.json before consuming this release. Its absence means the upload is incomplete. Pin the completed repository revision for experiments. Original, unpacked document text from uonlp/CulturaX. This is a fresh sample, independent of nguyenhuuthuat09/CulturaX_sampled. No token sequences or training caches are distributed. Each subset has a fixed validation set and a nested training prefix suitable for smaller token budgets. Token counts… See the full description on the dataset page: https://huggingface.co/datasets/nht10/cx_sampled.
CulturaX sampled text pools
Use UPLOAD_COMPLETE.json before consuming this release. Its absence means the upload is incomplete. Pin the completed repository revision for experiments.
Original, unpacked document text from uonlp/CulturaX. This is a fresh sample, independent of nguyenhuuthuat09/CulturaX_sampled. No token sequences or training caches are distributed. Each subset has a fixed validation set and a nested training prefix suitable for smaller token budgets.
Token counts use Qwen/Qwen3-0.6B-Base, revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd. Sampling seed: 42. Source revision: 6a8734bc69fefcbb7735f4f9250f43e4cd7a442e. Each count includes one explicit <|endoftext|> (151643) after the document; that marker is not appended to the stored text. Training must append it. Tokenization has no truncation or added special tokens. Budgets round up to whole documents and do not include packing-loss headroom. Other tokenizers can use the text, but these counts and token-budget anchors will not apply to them.
from datasets import load_dataset
data = load_dataset("nht10/cx_sampled", "qwen3_0.6b_base_ar_10B", revision="<completed-commit>")Each subsets/<name>/train/shard_* and validation/shard_* is also a complete save_to_disk folder: use load_from_disk and concatenate shards in sorted order. This preserves compatibility with the existing local training loader. Download all Arrow files referenced by each shard's state.json, plus state.json and dataset_info.json. document_index.parquet is a row-aligned sidecar, explicitly excluded from HF text split definitions.
subsets/<name>/manifest.json records ordered shard paths, counts, cumulative endpoints, checksums, source-file identities, and every 100M-token anchor. For an arbitrary smaller budget, select the prefix of training shards and use the boundary shard's cumulative token index to round up to a whole document. Keep its fixed validation split. The root manifest lists every published file. Individual files can be downloaded using pinned HF resolve URLs; a whole-repo download retrieves all ten subsets. Downloading files does not load a Dataset.
Sampling shuffles source files, row groups, and rows within groups with seeded RNGs; this is not a globally uniform document shuffle. Validation is selected first; its exact text hashes are excluded from training within each language. There is no additional global or near-duplicate deduplication. See provenance for the frozen sampling implementation, complete file choices and tokenizer.
Source and license
This sample inherits the source terms. CulturaX states that its licensing follows mC4 and OSCAR; see the preserved upstream README under provenance and the CulturaX dataset card. No new license or ownership of the underlying web documents is claimed here.
