the-malware-files/pentimento-core
Pentimento Core, v1.0.0 This is a JPEG-decompressed spatial corpus. It is NOT comparable to BOSSbase, whose covers were never JPEG compressed. Detector numbers measured here and numbers measured on BOSSbase cannot be placed in the same table. A steganalysis corpus of 10,000 permissively licensed cover photographs and 344,357 matched stego pairs across 35 arms, where every image carries its own licence and every file carries its own checksum. A further 4 arms hold the 38,119… See the full description on the dataset page: https://huggingface.co/datasets/the-malware-files/pentimento-core.
Pentimento Core, v1.0.0
This is a JPEG-decompressed spatial corpus. It is NOT comparable to BOSSbase, whose covers were never JPEG compressed. Detector numbers measured here and numbers measured on BOSSbase cannot be placed in the same table.
A steganalysis corpus of 10,000 permissively licensed cover photographs and 344,357 matched stego pairs across 35 arms, where every image carries its own licence and every file carries its own checksum. A further 4 arms hold the 38,119 clean halves those pairs are measured against.
Quick start
Nothing is downloaded until you ask for a sample, so looking costs seconds rather than 48 GB.
from datasets import load_dataset
covers = load_dataset("the-malware-files/pentimento-core", "covers", split="full", streaming=True)
print(next(iter(covers))["json"]["licence"])
stego = load_dataset("the-malware-files/pentimento-core", "wow-0050", split="full", streaming=True)Each arm is its own config, named exactly as the arm table below names it, and covers is the default. Drop streaming=True to fetch a config to disk.
The single split is called full because the shards are not laid out along the train and test boundary. Cover records carry their own split field and SPLITS.md gives the rule; splitting any other way puts a cover in training and its own stego copy in test.
What makes it different
Three things, each of which exists because something else was missing.
Adaptive schemes alongside real tools. REVEAL, from the Netherlands Forensic Institute, covers more than 50 end-user hiding tools and excludes content-adaptive academic schemes on purpose. This corpus carries HUGO, WOW, S-UNIWARD, HILL, MiPOD, J-UNIWARD and UERD as separate labelled arms.
Per-source arms, not a blend. The ranking of embedding schemes is known to invert when the cover source changes. A corpus that blends sources hides exactly the effect a user most needs to measure, so the arms here stay separate and labelled.
Rebuildable byte for byte. Every generator is seeded and every file is recorded with a sha256 and the count of samples actually changed. Results in the steganalysis literature differ by several accuracy points on the choice of data split alone, same network and same algorithm, so a corpus that cannot be rebuilt identically cannot support a comparison.
Layout
Shards are WebDataset tar files. Inside each, a sample's parts share a basename:
pentimento-core-00000.tar
000000.png the image
000000.json its manifest row, licence includedThey stream without unpacking, and every major dataset loader reads them.
The four digits in an arm name are not one quantity. hugo-0200 is 0.2 bits per pixel; juniward-0200 is 0.2 bits per non-zero AC coefficient; steghide-0200 is 0.2 of the capacity steghide itself reports. Ranking arms on the number in the name compares three different things. Every sample's JSON record carries rate_unit, which is the authoritative answer per file.
Every JPEG arm is written at quality 95, and the adaptive spatial arms are simulated at the optimal embedding rate rather than by a real STC coder - the coding field on each record says so. Both are ordinary practice and both change what a result means, so neither should be discovered after the fact.
Each part ships its own checksum file, SHA256SUMS-covers and SHA256SUMS-arms, so verifying a download is one command:
sha256sum -c SHA256SUMS-coversDo it before use: a shard that arrived truncated reads as a smaller corpus rather than as an error.
On Kaggle the shards are unpacked, and that command does not apply there. Kaggle extracts archives when they are uploaded and offers no way to refuse, so pentimento-core-00000.tar arrives as a folder of the same members under the same names. The bytes are the same; the container is gone. Verify that copy against each record's own checksum instead, which is a finer check because it names the file that is actually wrong:
python load_pentimento.py --verify pentimento-core-00000/load_pentimento.py reads a folder and a tar the same way, so nothing else changes. The Internet Archive and HuggingFace copies are tar shards as described above.
Tiers nest. Nano is the first 200 covers of the same ordering Lite's first 1,000 and Core's 10,000 follow, so you can develop against a small tier and evaluate on a larger one without the two overlapping in a way that flatters the result.
Licensing, in one paragraph
Every file carries its own licence, and each file's own `.json` member inside the shards is the authoritative record of it. ATTRIBUTION.csv beside this file is an extract of those records for the covers that require a credit line, and is the one to read if you want the licences without downloading the shards. (This paragraph used to point at "the manifest", which does not ship as a standalone file; a reader who wanted to check a licence before committing to 48 GB was sent to something they could not find.) The collection is published as CC BY 4.0, which is the strictest obligation present, not the loosest: complying with it satisfies every file here. Where a file's own record names something looser, rely on that instead.
5,453 of 10,000 covers (54.5%) require attribution, and each carries a ready-made credit line in its manifest row under attribution. Stego images are derivatives and inherit their cover's terms; every stego sample JSON carries the cover's licence under cover_licence, so a reader holding only one arm can still discharge the obligation.
See LICENCES.md for the full statement.
Before you train on it
Read SPLITS.md. A cover and its stego versions are far more alike than any two unrelated photographs, so a random split puts a cover in training and its own stego copy in test, and the classifier learns the photograph rather than the payload. Split by cover, never by image.
Citing
See CITATION.cff, which GitHub and Zenodo both render automatically.
Provenance
Covers come from Wikimedia Commons under permissive licences, deduplicated by perceptual hash across sources and sessions, and capped per photographer and per camera body so no single prolific uploader dominates the sample. Rejected candidates are recorded with their reason rather than silently dropped.
A caveat on that last sentence, kept because it was not true once. Cover SELECTION records its rejections. The arm BUILDERS did not: a failure during embedding was written to stderr and the loop continued, so 118 covers were missing from every spatial arm with nothing in the corpus saying so, and the build log that held the reason had been superseded. The images turned out to have been written without their manifest rows, they were rebuilt, and the arms are complete. The builders now record what they skip.
