AyoubChLin/Company-document-dataset-v2
Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.
Company Documents v2
Generation complete: all 13 document types have completed export and upload checkpoints.
Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types).
Dataset overview
Intended uses
Document classification, OCR evaluation, key information extraction, layout-aware field extraction, and document question answering using the gold fields to construct task-specific examples. The dataset provides document text, structured fields and coordinates; it does not include a separate set of question-answer pairs.
Subsets
Pick a document type, or all. Every subset has the same train / validation / test split.
Quick start
import json
from datasets import load_dataset
documents = load_dataset("AyoubChLin/Company-document-dataset-v2", "invoice", split="train")
row = documents[0]
row["pdf"] # pdfplumber.PDF
json.loads(row["extracted_data"]) # the gold fieldsUse "all" to load every document type. For a large subset, pass streaming=True to load_dataset and read examples with next(iter(documents)) to avoid downloading the entire subset.
Scans
120,523 degraded page images (scan and photo profiles) of a sample of the documents, with the gold-field boxes moved into image pixels:
scans = load_dataset("AyoubChLin/Company-document-dataset-v2", "scans", split="test")
scans[0]["image"], json.loads(scans[0]["boxes"])Files
data/<type>/<split>-NNNNN.parquet one row per PDF, the PDF embedded (the columns below)
scans/<type>/<split>-NNNNN.parquet one row per degraded page image
stats.json counts per split, type, source, locale, layout; companies per split
licenses/ source database licensesSplits
Splits are by issuing company (per sector): validation and test documents come from companies whose issuing company is absent from training. Source databases, products, counterparties, layout families and themes can occur across splits; this is an issuer-held-out split, not a source-held-out split.
Columns
One row per PDF (all and every document-type subset):
scans subset, one row per degraded page image (profiles scan and photo: skew, blur, noise, paper tint, lighting, JPEG):
Sources
Northwind (food wholesale), AdventureWorks (bicycle manufacturer: sales, purchasing, production, HR), Chinook (online music store), Sakila (video rental, dates shifted +19 years). Issuers are synthetic; counterparties, products, quantities and prices come from the databases.
Guarantees
Every document passed a self-check: the checked gold values are present in the PDF text and the arithmetic holds (line totals, tax, totals, running balances, aging, payroll, received vs accepted). Money is computed with Decimal and half-up rounding; currencies are converted once from the source (USD).
Limitations
- These are synthetic documents derived from sample databases. Their layouts, wording and simulated degradation do not cover the full variety of real business documents or camera captures.
- Document types are imbalanced; use the per-type counts when selecting training data and reporting evaluation results. Some types contain only a small number of examples.
- Scan rows represent pages of sampled PDFs, not additional independent business transactions. Keep scans with their parent document's split when constructing an evaluation dataset.
- Gold fields and word boxes come from the generation and PDF extraction pipeline. Arithmetic and text checks do not establish that every annotation or reading-order decision is error-free.
- The same underlying source records can support multiple document types. Holding out issuing companies does not guarantee that all business entities or transaction content are unseen.
Licenses
Code and generated documents: Apache-2.0. Source data: Northwind (MIT), AdventureWorks (MIT), Chinook (MIT), Sakila (BSD-2); see licenses/.
Citation
Created by Cherguelaine Ayoub. If you use this dataset, please cite:
@misc{cherguelaine2026companydocumentsv2,
author = {Cherguelaine, Ayoub},
title = {Company Documents v2: Synthetic Business Documents},
year = {2026},
url = {https://github.com/AyoubCherguelaine/Company-document-dataset-v2}
}Reproduce
Generator source code: Company-document-dataset-v2. To generate the content locally from a Git checkout:
git clone https://github.com/AyoubCherguelaine/Company-document-dataset-v2.git
cd Company-document-dataset-v2
# Generate all source records and their scans locally.
python3 scripts/run_pipeline.py --all-records --stages venv download index companies texts generate augment --seed 42 --workers 8 --augment-fraction 0.3
# Export with the published dataset's company split seed.
./venv/bin/python -m docgen export output/v2 dataset/full-local --format parquet --seed 0See the repository README for setup, configuration and full-dataset generation options.
Generated with the CompanyDocuments v2 generator (docgen) and this run config:
{
"run_name": "full",
"seed": 42,
"out_dir": "output/parts/<type>",
"workers": 8,
"fx": {
"USD": "1",
"EUR": "0.92",
"GBP": "0.79"
},
"self_check": true,
"boxes": true,
"resume": true,
"documents": [
{
"type": "account_statement",
"count": "all"
},
{
"type": "credit_note",
"count": "all"
},
{
"type": "employment_certificate",
"count": "all"
},
{
"type": "goods_received_note",
"count": "all"
},
{
"type": "inventory_report",
"count": "all"
},
{
"type": "invoice",
"count": "all"
},
{
"type": "packing_slip",
"count": "all"
},
{
"type": "payslip",
"count": "all"
},
{
"type": "purchase_order",
"count": "all"
},
{
"type": "quote",
"count": "all"
},
{
"type": "receipt",
"count": "all"
},
{
"type": "shipping_order",
"count": "all"
},
{
"type": "work_order",
"count": "all"
}
]
}