StentorLabs/StenCore-PDF
StenCore — FinePDFs-Edu Curated By StentorLabs StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting. ⚠️… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.
StenCore — FinePDFs-Edu Curated
By [StentorLabs](https://huggingface.co/StentorLabs)
StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting.
⚠️ Privacy & Copyright Notice: This dataset was produced by automated pipelines that include regex-based PII redaction and heuristic filters. Automated redaction may miss personal data (false negatives) and may over-redact (false positives). No claim of complete PII removal is made. Do not use this dataset in systems that surface individual personal data. Review and clear upstream site-level licenses before redistributing or deploying models trained on this data. For takedown or privacy requests, open an issue or contact stentorlabs@gmail.com with the affected doc identifier and evidence.
Quick Stats
Usage
from datasets import load_dataset
ds = load_dataset("StentorLabs/stencore")
# Filter to highest-quality documents
top = ds["train"].filter(lambda r: r["cqf_score"] >= 0.5)
print(len(top))
print(top[0]["text"][:500])Intended & Disallowed Uses
✅ Intended uses
- LLM pre-training and fine-tuning on English educational text
- Quality filtering and data curation research
- Domain reweighting and curriculum learning experiments
🚫 Disallowed uses
- Systems that surface or process personal information
- Redistribution without reviewing upstream licensing (Apache 2.0 requires attribution; check site-level licenses for individual documents)
- Any production deployment without independent legal and privacy review
Pipeline Summary
StenCore (curate_2026 mode) is a fully automated curation pipeline (v2026.03) running 14 sequential stages from raw Parquet to a publish-ready HuggingFace dataset.
Every document in this dataset passed all of the following:
- Columnar pre-filter — DuckDB SQL-level filter over raw Parquet
- Adaptive quality thresholds — learned from the data, not hand-tuned (per-register: formal / conversational / technical)
- Language ID gate — fastText or stopword/script heuristic (
enonly) - Heuristic quality filter — word count, alpha ratio, stopword hits, punctuation ratio, repetition ratios, avg word length
- PII redaction — regex detection of emails, phones, SSNs, IPs, card numbers, IBANs, API keys
- Toxicity screening — lexical axis scoring, drop policy
- Eval decontamination — exact, 8-gram overlap, char 3-gram, SimHash checks against benchmark sets
- Exact deduplication — SHA-1 fingerprint (in-memory)
- MinHash near-dedup — token shingle LSH (autotuned to 25% sample rate)
- CQF quality gate — top 85% by quality score (threshold: CQF ≥ 0.4167)
- KenLM + neural perplexity gate —
edugp/kenlmWikipedia model + SmolLM2-135M (excess mode, ref mean: 916.24) - Cluster rehydration — restores high-quality representatives from dedup-dropped clusters
- Domain reweighting — per-domain weights learned from proxy model evaluation
- Mix optimization + final write — domain-quota sampling, curriculum ordering, shard + merkle output
<details> <summary><b>📋 Full Stage-by-Stage Breakdown</b></summary>
Pre-Stage: Environment & Model Setup
KenLM Scorer
- Repo:
edugp/kenlm| Corpus:wikipedia/en - Snapshot:
3fbe35c83b1a39f420a345b7c96a186c8030d834 - Mode:
first_pass— KenLM prefilters, SmolLM2 rescores the subset
Neural Reference LM
- Model:
HuggingFaceTB/SmolLM2-135M(torchaoint8_weight_only) torch.compile: Disabled (CPU stability)
Effective runtime profile:
light_mode=False input_cap=400 bootstrap_docs=400
stage_timeout_s=90 stream_chunk=32768 ppl_batch=192
ppl_workers=1 prefix_tokens=256 kenlm_mode=first_pass12h target profile:
hf_input_target_gb=7.50 hf_input_max_files=64
cqf_keep_web=0.85 prior_keep=0.90 web_ratio=0.95
strict_ratio=False dedup=memory autotune_grace=10000Stage 1: resolvewebinputs
Wall: 53.47s | CPU: 143.36s | Ratio: 2.68×
DuckDB columnar prefilter over 3 Parquet files → stage_columnar_prefilter.parquet. High CPU/wall ratio confirms parallel DuckDB execution.
Stage 2: adaptmultilingualprofiles
Wall: 25.48s | CPU: 25.39s | Ratio: 1.00×
Draws a pilot sample and learns per-register quality threshold parameters from the data. Three registers profiled: en:formal (1,854 samples), en:conversational (2,993 samples), en:technical. Single-threaded statistical pass.
Stage 3: optimizethresholdprofiles
Wall: 86.79s | CPU: 138.40s | Ratio: 1.59×
Grid search over code/math profile scales, ranked by proxy quality metric. Selects best-performing threshold configuration.
Stage 4: cqfseedverification_loop
Wall: 0.001s | CPU: 0.001s
Retrains CQF fastText at multiple thresholds, picks best by proxy quality. Completed near-instantly — seed was valid, no remediation needed.
Stage 5: stagewebcandidate_pass
Wall: 27,162.57s (7.5h) | CPU: 42,229.32s | Ratio: 1.55×
Main filtering stage. Sequential gates per document:
- Quick prefilter (min words / alpha)
- HF-specific input filters (English subset, EDU v2, strip code fences)
- Source policy (host allow/deny, ToS risk, license allowlist)
- Line cleaner (boilerplate, nav, HTML artifacts, duplicate lines)
- Domain routing (code / math / prose)
- Language ID gate (+ English noise guard, code bypass)
- Domain-aware heuristic filter (adaptive thresholds from Stage 2/3)
- PII redaction
- Toxicity screening
- Eval decontamination (exact, 8-gram, char 3-gram, SimHash)
- CQF scoring
- Exact dedup (SHA-1, in-memory)
- MinHash near-dedup (autotuned)
- Semantic dedup (autotuned to
none)
Kept: 220,635 / 584,000 (37.78%). Dropped records to semantic dup pool for potential rehydration.
Stage 6 (observed): midproxystage3
Wall: 115.50s | CPU: 160.80s | Ratio: 1.39×
Intermediate proxy scoring pass over quality-filtered candidates.
Stage 7: stagewebqualityandperplexity
Wall: 300.05s | CPU: 282.49s | Ratio: 0.94×
- CQF threshold gate — keeps top 85% (CQF ≥ 0.4167)
- Multi-objective property minima — per-property floor checks
- Prior noise gate — filters statistical outliers vs. quality prior
- Hybrid disagreement trigger — routes CQF/secondary disagreements to full neural perplexity
- KenLM + SmolLM2 perplexity gate — excess mode, ref mean 916.24
22,450 documents scored. Avg perplexity: 902.02 (stored as excess perplexity — relative to the 916.24 calibrated mean, not absolute).
Stage 8: rehydrate_clusters
Wall: 40.42s | CPU: 40.19s | Ratio: 0.99×
Reads the semantic dup pool. Re-adds top-quality cluster representatives (ranked by FineWeb2-like weighted formula). Optional MMR diversity selection. Single-threaded.
Stage 9 (observed): midproxystage5
Wall: 102.92s | CPU: 143.10s | Ratio: 1.39×
Second intermediate proxy scoring pass after cluster rehydration.
Proxy Eval & Domain Reweighting (~60 min observed gap)
Per-domain proxy scores aggregated → per-domain reweighting coefficients learned. Hundreds of source domains reweighted. Examples:
Stage 9: buildsyntheticpool
Wall: 0.003s — No teacher LLM configured. Pool: 0 documents.
Stage 10: stagesyntheticfilter
Wall: 0.002s — No-op (empty pool).
Stage 11: mix_optimization
Wall: 175.34s | CPU: 214.87s | Ratio: 1.23×
Proxy-evaluated search over web/synth mixing ratios. Applies domain weights from proxy eval. Produces final document selection plan.
Stage 12: finalmixand_write
Wall: ~2,311s (~38.5 min)
Domain quota enforcement → domain-weighted acceptance sampling → final exact dedup → optional curriculum ordering (easy→hard by CQF/perplexity) → streaming JSONL + Parquet write → SHA-256 + merkle root → optional deterministic sharding.
Stages 13–14: proxyeval + hfpush
Final proxy metric gate enforcement → auto-upload to HuggingFace Hub.
</details>
Adaptive Quality Thresholds
<details> <summary><b>📐 Full learned threshold values (en:formal and en:conversational)</b></summary>
These parameters were learned from the data in Stage 2 and optimized in Stage 3. All values are exact as logged.
en:formal — 1,854 bootstrap samples
en:conversational — 2,993 bootstrap samples
en:technical
Separate profile applied; full values truncated in logs. Expected to have higher tolerance for non-alphabetic characters (equations, code, symbols) and relaxed stopword requirements.
</details>
Deduplication Autotune
<details> <summary><b>⚙️ Runtime autotune event log</b></summary>
The autotune system fires after 10,000 documents (grace period). All 6 events occurred in a 5,000-document window — the system converges aggressively.
Throughput gain: ~1.8× (12.3 → 22 docs/s). MinHash at 25% sampling will miss some near-duplicate pairs — a deliberate throughput/precision tradeoff accepted by the autotune system based on observed duplicate density.
</details>
Perplexity Scoring
<details> <summary><b>📊 KenLM + SmolLM2 scoring details</b></summary>
Configuration
- KenLM model:
edugp/kenlm,wikipedia/en, snapshot3fbe35c83b1a39f420a345b7c96a186c8030d834 - Neural LM:
HuggingFaceTB/SmolLM2-135M(torchao int8weightonly) - Mode:
first_pass— KenLM prefilters all docs; SmolLM2 rescores near-boundary subset - Scoring mode:
excess—score = raw_perplexity - reference_mean - Reference mean: 916.2404 (calibrated fresh on 256 bootstrap docs,
fit_new)
Scoring progress
High individual perplexity values (>7,000) are expected for math-heavy or notation-rich educational text under a Wikipedia-trained model. The excess mode partially normalizes this. Running mean converges to ~902 across the full scored set.
</details>
Runtime & Timing
<details> <summary><b>⏱️ Full per-stage timing table</b></summary>
Candidate pass throughput: ~12.3 docs/s (initial) → ~22 docs/s (post-autotune, ~1.8× gain).
</details>
Limitations
- PDF extraction artifacts — OCR artifacts, broken equations, and malformed tables may be present despite filtering.
- Residual PII — Automated regex redaction does not guarantee complete PII removal. Do not use for systems that surface personal information.
- Copyright — Source PDFs may carry individual site-level licenses. Apache 2.0 requires attribution; verify upstream licensing for your use case.
- KenLM Wikipedia bias — Math-heavy or highly technical documents may be underrepresented due to high perplexity under a Wikipedia-trained model.
- ~62% rejection rate — Some valid educational content may have been dropped due to heuristic threshold mismatch (e.g., table-heavy or equation-dense formatting).
- English only — Pipeline profiled
en:formal,en:conversational, anden:technicalregisters only. - No synthetic data — This run did not use the synthetic generation system. Dataset is 100% source text.
- MinHash 25% sampling — Post-autotune dedup will miss some near-duplicate pairs.
Licensing & Citation
Released under Apache 2.0. Attribution required. Derived from HuggingFaceFW/finepdfs-edu — review upstream licensing before use.
@dataset{stencore_finepdfs_edu_curated,
title = {StenCore: FinePDFs-Edu Curated},
author = {StentorLabs},
year = {2026},
note = {StentorLabs' first dataset. StenCore pipeline v2026.03.
584k docs in, 149k out. Adaptive heuristics, PII redaction,
toxicity/decontam gates, MinHash + KenLM + neural perplexity,
CQF scoring, proxy domain reweighting.},
howpublished = {\url{https://huggingface.co/datasets/StentorLabs/stencore}}
}
@dataset{fineweb_finepdfs_edu,
author = {HuggingFace FineWeb Team},
title = {FinePDFs-Edu},
howpublished = {\url{https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu}}
}Contact: StentorLabs@gmail.com — for takedown requests, privacy concerns, or feedback.
<p align="center">Made with ❤️ by <a href="https://huggingface.co/StentorLabs">StentorLabs</a></p>
