datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.StenCore-PDF
StenCore — FinePDFs-Edu Curated
By StentorLabs
StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting.
⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.pdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.napierone-pdf-raw
BEE-spoke-data/napierone-pdf-raw
NapierOne PDF files converted with marker.
detected languages
Counter({'en': 4665,
'nl': 2,
'fi': 7,
'fr': 8,
'cy': 54,
'sq': 1,
'it': 1,
'unknown-error': 5,
'sk': 1,
'es': 2,
'de': 3,
'ro': 1,
'pl': 1,
'zh': 1,
'so': 1,
'ml': 1})
