Team Ai
Datasetpublic

AyoubChLin/Company-document-dataset-v2

Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
2likes12kdownloads
Dataset Card

Company Documents v2

Generation complete: all 13 document types have completed export and upload checkpoints.

Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types).

Dataset overview

propertyvalue
PDF documents353,580
PDF pages404,514
Document types13
Synthetic issuing companies60
LanguagesEnglish / French
Degraded scan/photo page images120,523
Export formatparquet
Split unitissuing company, grouped by sector

Intended uses

Document classification, OCR evaluation, key information extraction, layout-aware field extraction, and document question answering using the gold fields to construct task-specific examples. The dataset provides document text, structured fields and coordinates; it does not include a separate set of question-answer pairs.

Subsets

Pick a document type, or all. Every subset has the same train / validation / test split.

subsetdocumentdocumentstrainvalidationtestsources
all (default)every type below353,580259,49747,11446,969all four
account_statementa customer's invoices and payments over a period, running balance and aging72852299107adventureworks, northwind
credit_notecredit against an invoice, with a reason (damaged, wrong item, price...)47,74735,1076,3376,303adventureworks, northwind
employment_certificateletter certifying an employee's job title, department and hire date2902024345adventureworks
goods_received_notedelivery check against a purchase order: ordered, received, accepted, rejected3,6892,667512510adventureworks
inventory_reportstock on hand per product, with value and reorder status5135610adventureworks, northwind
invoicesales invoice: line items, discounts, tax, totals, payment terms48,15835,3426,4466,370adventureworks, chinook, northwind
packing_slipitems packed for a shipment: ordered vs shipped quantities47,74035,0406,4566,244adventureworks, northwind
payslipemployee pay for a period: earnings, deductions, net pay2902313029adventureworks
purchase_orderorder the issuer places with a vendor50,31336,9086,5986,807adventureworks, northwind
quoteprice quotation with validity and lead time47,74635,0046,3136,429adventureworks, northwind
receiptproof of a paid retail transaction (A5)16,45612,1052,2182,133chinook, sakila
shipping_orderinstruction to ship an order: items, weights, consignee, carrier47,74735,1826,2746,291adventureworks, northwind
work_orderproduction order: routing operations, planned vs actual cost42,62531,1525,7825,691adventureworks
scansdegraded page images of a sample of the documents120,523 pages88,79915,82915,895

Quick start

python
import json
from datasets import load_dataset

documents = load_dataset("AyoubChLin/Company-document-dataset-v2", "invoice", split="train")
row = documents[0]
row["pdf"]                            # pdfplumber.PDF
json.loads(row["extracted_data"])     # the gold fields

Use "all" to load every document type. For a large subset, pass streaming=True to load_dataset and read examples with next(iter(documents)) to avoid downloading the entire subset.

Scans

120,523 degraded page images (scan and photo profiles) of a sample of the documents, with the gold-field boxes moved into image pixels:

python
scans = load_dataset("AyoubChLin/Company-document-dataset-v2", "scans", split="test")
scans[0]["image"], json.loads(scans[0]["boxes"])

Files

data/<type>/<split>-NNNNN.parquet    one row per PDF, the PDF embedded (the columns below)
scans/<type>/<split>-NNNNN.parquet   one row per degraded page image
stats.json                           counts per split, type, source, locale, layout; companies per split
licenses/                            source database licenses

Splits

Splits are by issuing company (per sector): validation and test documents come from companies whose issuing company is absent from training. Source databases, products, counterparties, layout families and themes can occur across splits; this is an issuer-held-out split, not a source-held-out split.

splitdocumentscompaniesscan pages
train259,4974488,799
validation47,114815,829
test46,969815,895

Columns

One row per PDF (all and every document-type subset):

columncontent
pdfthe PDF itself (the viewer shows its first page; datasets opens it with pdfplumber; raw bytes with .cast_column("pdf", Pdf(decode=False)))
file_name<type>/<doc_id>.pdf
doc_id, split, document_typeidentifiers
source, variantsource database and document variant (<type>.<source>)
company, company_name, company_sector, company_country, company_city, company_tax_id, company_locale, company_currencythe issuing company
number, issue_date, currency, total, items_count, counterparty_role, counterparty_namekey gold values as plain columns
layout, theme, locale, pagesletterhead layout, color theme, language, page count
extracted_dataall gold fields as JSON
file_contenttext extracted from the PDF, in reading order
optionscontent variation drawn for the document (JSON)
renderedeach gold field as printed in the PDF (JSON)
boxeswhere each gold field is printed (JSON: field path -> list of page + bbox in PDF points)
wordsevery word of the PDF with its page and bbox (JSON)

scans subset, one row per degraded page image (profiles scan and photo: skew, blur, noise, paper tint, lighting, JPEG):

columncontent
imagethe page image
doc_id, split, document_type, page, profileidentifiers; join to the documents on doc_id
width, heightimage size in pixels
paramsdegradation parameters (JSON)
words, boxeswords and gold-field boxes moved into image pixels (JSON)

Sources

sourcedocuments
adventureworks208,870
northwind127,842
sakila16,044
chinook824

Northwind (food wholesale), AdventureWorks (bicycle manufacturer: sales, purchasing, production, HR), Chinook (online music store), Sakila (video rental, dates shifted +19 years). Issuers are synthetic; counterparties, products, quantities and prices come from the databases.

languagedocuments
English203,657
French149,923

Guarantees

Every document passed a self-check: the checked gold values are present in the PDF text and the arithmetic holds (line totals, tax, totals, running balances, aging, payroll, received vs accepted). Money is computed with Decimal and half-up rounding; currencies are converted once from the source (USD).

Limitations

  • —These are synthetic documents derived from sample databases. Their layouts, wording and simulated degradation do not cover the full variety of real business documents or camera captures.
  • —Document types are imbalanced; use the per-type counts when selecting training data and reporting evaluation results. Some types contain only a small number of examples.
  • —Scan rows represent pages of sampled PDFs, not additional independent business transactions. Keep scans with their parent document's split when constructing an evaluation dataset.
  • —Gold fields and word boxes come from the generation and PDF extraction pipeline. Arithmetic and text checks do not establish that every annotation or reading-order decision is error-free.
  • —The same underlying source records can support multiple document types. Holding out issuing companies does not guarantee that all business entities or transaction content are unseen.

Licenses

Code and generated documents: Apache-2.0. Source data: Northwind (MIT), AdventureWorks (MIT), Chinook (MIT), Sakila (BSD-2); see licenses/.

Citation

Created by Cherguelaine Ayoub. If you use this dataset, please cite:

bibtex
@misc{cherguelaine2026companydocumentsv2,
  author = {Cherguelaine, Ayoub},
  title = {Company Documents v2: Synthetic Business Documents},
  year = {2026},
  url = {https://github.com/AyoubCherguelaine/Company-document-dataset-v2}
}

Reproduce

Generator source code: Company-document-dataset-v2. To generate the content locally from a Git checkout:

bash
git clone https://github.com/AyoubCherguelaine/Company-document-dataset-v2.git
cd Company-document-dataset-v2
# Generate all source records and their scans locally.
python3 scripts/run_pipeline.py --all-records --stages venv download index companies texts generate augment --seed 42 --workers 8 --augment-fraction 0.3
# Export with the published dataset's company split seed.
./venv/bin/python -m docgen export output/v2 dataset/full-local --format parquet --seed 0

See the repository README for setup, configuration and full-dataset generation options.

Generated with the CompanyDocuments v2 generator (docgen) and this run config:

yaml
{
  "run_name": "full",
  "seed": 42,
  "out_dir": "output/parts/<type>",
  "workers": 8,
  "fx": {
    "USD": "1",
    "EUR": "0.92",
    "GBP": "0.79"
  },
  "self_check": true,
  "boxes": true,
  "resume": true,
  "documents": [
    {
      "type": "account_statement",
      "count": "all"
    },
    {
      "type": "credit_note",
      "count": "all"
    },
    {
      "type": "employment_certificate",
      "count": "all"
    },
    {
      "type": "goods_received_note",
      "count": "all"
    },
    {
      "type": "inventory_report",
      "count": "all"
    },
    {
      "type": "invoice",
      "count": "all"
    },
    {
      "type": "packing_slip",
      "count": "all"
    },
    {
      "type": "payslip",
      "count": "all"
    },
    {
      "type": "purchase_order",
      "count": "all"
    },
    {
      "type": "quote",
      "count": "all"
    },
    {
      "type": "receipt",
      "count": "all"
    },
    {
      "type": "shipping_order",
      "count": "all"
    },
    {
      "type": "work_order",
      "count": "all"
    }
  ]
}