akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this exists
Frontier models have read the public internet. They have not read the operational record of how physical trade actually runs: the back-and-forth between a forwarder and a customs broker, the rate that got negotiated, the wagon that went missing, the invoice that was reconciled three weeks later.
That record exists, but it sits inside companies and is never indexed. This is a slice of one such archive, cleared for release.
Contents
Languages, by document
The corpus is predominantly Russian. It reflects how business in the region actually communicated in this period, including code-switching between Russian and Ukrainian inside a single thread — which is authentic, and is not something a curated single-language corpus will give you.
How the text was obtained
Three quarters of this corpus existed only as scanned paper. It was recovered by OCR in Ukrainian and Russian at a mean per-word confidence of 90.23. Every OCR document carries an ocr object recording engine, model, languages and confidence, plus a warnings entry noting that the text may contain recognition errors. Filter on that field if your use requires faithful text.
What is in it
Commercial rail freight correspondence and the documents attached to it: SMGS consignment notes, acts of completed services, invoices, wagon manifests and appendices. Cargo covered includes coal, alumina, grain, oilseed meal, petroleum products, timber and construction aggregates, moving across CIS routes.
Commercial content is intact by design. Wagon numbers, consignment and invoice numbers, station and route names, tonnages, tariffs, monetary amounts and contract dates all survive de-identification. That is what makes the corpus worth training on, and stripping it would have left prose about nothing.
What was removed
Direct identifiers were replaced with stable surrogates. The same source value always yields the same label, so threading, co-occurrence and relational structure survive.
Record identifiers, content hashes, storage paths and RFC 5322 Message-IDs were re-keyed under HMAC with a run-specific salt, so they do not match the source archive.
The re-identification key was destroyed. The key map and the HMAC salt were deleted after generation and are not retained by the publisher. No reasonably likely means of re-identification remains from this release alone (GDPR Recital 26). This release is anonymous rather than pseudonymised.
A residual-identifier gate was run over the output before release: passed.
Original .eml files and attachment binaries are not part of this release.
Limitations you should know about
- De-identification is rule-based.
DEID_METHOD.mddocuments the method and states plainly what it can and cannot catch. - OCR text carries recognition errors. Mean confidence is 90.23, not 100.
- Timestamps are at second precision. Against a small correspondent set these are a quasi-identifier for anyone holding correlating records.
- Company names are retained. Legal entities are not personal data under GDPR, but a small company plus a role can single out an individual. This was a deliberate trade-off in favour of commercial utility.
- One residual company code survives: a bank's own public registration number appears in account-details boilerplate, because the keyword rule expected the identifier to follow the label directly and an intervening word broke the match. It identifies a large public bank, not a person or the originating company.
- 90 messages is a sample, not a corpus. It is sized for evaluation.
Files
Generated by pipeline deid-1.0.
Suggested uses
- Pre-training and continued pre-training on non-English business language
- Document understanding over real, messy, scanned commercial paperwork
- Information extraction: entities, amounts, references, routes
- Agentic and workflow-reasoning evaluation — the records link into complete operational sequences, from enquiry through rate calculation, contract, invoice, dispatch, tracking, customs and reconciliation
- Low-resource and code-switching language work in Ukrainian and Russian
Provenance and rights
The source is a Ukrainian rail-freight forwarding company. It is not named here, by agreement. Rights to the archive are held under written agreement with the originating business, and provenance documentation and the licence chain are available under NDA to serious counterparties.
Full archive
This sample is drawn from roughly 1.2 TB of operational records spanning 2014–2022 — email, contracts, transport and customs documents, spreadsheets and transaction records, across import, export and transit movements in the CIS, the Baltic states and Bulgaria. Further volume is delivered de-identified and normalised to this same schema, scoped to buyer requirements.
Akuma London sources private operational archives and builds them into model-ready corpora. We also hold further regional company archives and commission first-party egocentric video capture inside working industrial environments.
Enquiries: contact@akumalondon.com · akumalondon.com
Citation
@misc{akumalondon_rail_freight_2026,
title = {Ukrainian Rail-Freight Correspondence Corpus (Sample)},
author = {Akuma London},
year = {2026},
url = {https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample}
}Licence
Released under CC BY 4.0. You may use this sample commercially, including for model training, provided you give attribution to Akuma London.
