Team Ai
Datasetpublic

akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample

Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.

sourceHugging Facecc-by-4.0updated 28d agoView on Hugging Face
0likes108downloads
Dataset Card

Ukrainian Rail-Freight Correspondence Corpus (Sample)

Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before.

This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below.

Published by Akuma London · akumalondon.com

Why this exists

Frontier models have read the public internet. They have not read the operational record of how physical trade actually runs: the back-and-forth between a forwarder and a customs broker, the rate that got negotiated, the wagon that went missing, the invoice that was reconciled three weeks later.

That record exists, but it sits inside companies and is never indexed. This is a slice of one such archive, cleared for release.

Contents

Email messages90
Text documents (bodies + attachment units)1,076
Pre-training chunks1,126
Characters of text2,322,106
Date range2022-01-10 – 2022-07-26

Languages, by document

LanguageDocuments
Russian869
Ukrainian141
English / Latin62
Mixed or undetermined4

The corpus is predominantly Russian. It reflects how business in the region actually communicated in this period, including code-switching between Russian and Ukrainian inside a single thread — which is authentic, and is not something a curated single-language corpus will give you.

How the text was obtained

Extraction statusDocuments
extracted (native text)259
extracted_ocr (recovered from scans)817

Three quarters of this corpus existed only as scanned paper. It was recovered by OCR in Ukrainian and Russian at a mean per-word confidence of 90.23. Every OCR document carries an ocr object recording engine, model, languages and confidence, plus a warnings entry noting that the text may contain recognition errors. Filter on that field if your use requires faithful text.

What is in it

Commercial rail freight correspondence and the documents attached to it: SMGS consignment notes, acts of completed services, invoices, wagon manifests and appendices. Cargo covered includes coal, alumina, grain, oilseed meal, petroleum products, timber and construction aggregates, moving across CIS routes.

Commercial content is intact by design. Wagon numbers, consignment and invoice numbers, station and route names, tonnages, tariffs, monetary amounts and contract dates all survive de-identification. That is what makes the corpus worth training on, and stripping it would have left prose about nothing.

What was removed

Direct identifiers were replaced with stable surrogates. The same source value always yields the same label, so threading, co-occurrence and relational structure survive.

ClassDistinct valuesOccurrences replaced
PERSON441,265
EMAIL241,630
PHONE16125
EDRPOU (company registration no.)9574
IBAN759
TAXID659
MFO (bank code)542

Record identifiers, content hashes, storage paths and RFC 5322 Message-IDs were re-keyed under HMAC with a run-specific salt, so they do not match the source archive.

The re-identification key was destroyed. The key map and the HMAC salt were deleted after generation and are not retained by the publisher. No reasonably likely means of re-identification remains from this release alone (GDPR Recital 26). This release is anonymous rather than pseudonymised.

A residual-identifier gate was run over the output before release: passed.

Original .eml files and attachment binaries are not part of this release.

Limitations you should know about

  • —De-identification is rule-based. DEID_METHOD.md documents the method and states plainly what it can and cannot catch.
  • —OCR text carries recognition errors. Mean confidence is 90.23, not 100.
  • —Timestamps are at second precision. Against a small correspondent set these are a quasi-identifier for anyone holding correlating records.
  • —Company names are retained. Legal entities are not personal data under GDPR, but a small company plus a role can single out an individual. This was a deliberate trade-off in favour of commercial utility.
  • —One residual company code survives: a bank's own public registration number appears in account-details boilerplate, because the keyword rule expected the identifier to follow the label directly and an intervening word broke the match. It identifies a large public bank, not a person or the originating company.
  • —90 messages is a sample, not a corpus. It is sized for evaluation.

Files

PathContents
corpus/train.jsonlpre-training chunks; exact substrings of normalized/documents.jsonl
normalized/documents.jsonlfull text per document, with extraction provenance
normalized/messages.jsonlmessage headers, threading and attachment links
indexes/attachments.jsonlattachment-to-message relationships
indexes/source_files.jsonlsource file inventory
DATA_CARD.mdgenerated dataset card
DEID_METHOD.mdde-identification method statement
reports/DEID_REPORT.mdwhat was replaced, and the residual-identifier gate
reports/ocr_report.jsonper-file OCR coverage and confidence

Generated by pipeline deid-1.0.

Suggested uses

  • —Pre-training and continued pre-training on non-English business language
  • —Document understanding over real, messy, scanned commercial paperwork
  • —Information extraction: entities, amounts, references, routes
  • —Agentic and workflow-reasoning evaluation — the records link into complete operational sequences, from enquiry through rate calculation, contract, invoice, dispatch, tracking, customs and reconciliation
  • —Low-resource and code-switching language work in Ukrainian and Russian

Provenance and rights

The source is a Ukrainian rail-freight forwarding company. It is not named here, by agreement. Rights to the archive are held under written agreement with the originating business, and provenance documentation and the licence chain are available under NDA to serious counterparties.

Full archive

This sample is drawn from roughly 1.2 TB of operational records spanning 2014–2022 — email, contracts, transport and customs documents, spreadsheets and transaction records, across import, export and transit movements in the CIS, the Baltic states and Bulgaria. Further volume is delivered de-identified and normalised to this same schema, scoped to buyer requirements.

Akuma London sources private operational archives and builds them into model-ready corpora. We also hold further regional company archives and commission first-party egocentric video capture inside working industrial environments.

Enquiries: contact@akumalondon.com · akumalondon.com

Citation

bibtex
@misc{akumalondon_rail_freight_2026,
  title  = {Ukrainian Rail-Freight Correspondence Corpus (Sample)},
  author = {Akuma London},
  year   = {2026},
  url    = {https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample}
}

Licence

Released under CC BY 4.0. You may use this sample commercially, including for model training, provided you give attribution to Akuma London.