Team Ai
Datasetpublic

the-malware-files/pentimento-core

Pentimento Core, v1.0.0 This is a JPEG-decompressed spatial corpus. It is NOT comparable to BOSSbase, whose covers were never JPEG compressed. Detector numbers measured here and numbers measured on BOSSbase cannot be placed in the same table. A steganalysis corpus of 10,000 permissively licensed cover photographs and 344,357 matched stego pairs across 35 arms, where every image carries its own licence and every file carries its own checksum. A further 4 arms hold the 38,119… See the full description on the dataset page: https://huggingface.co/datasets/the-malware-files/pentimento-core.

sourceHugging Facecc-by-4.0updated 15d agoView on Hugging Face
0likes2.4kdownloads
Dataset Card

Pentimento Core, v1.0.0

This is a JPEG-decompressed spatial corpus. It is NOT comparable to BOSSbase, whose covers were never JPEG compressed. Detector numbers measured here and numbers measured on BOSSbase cannot be placed in the same table.

A steganalysis corpus of 10,000 permissively licensed cover photographs and 344,357 matched stego pairs across 35 arms, where every image carries its own licence and every file carries its own checksum. A further 4 arms hold the 38,119 clean halves those pairs are measured against.

Quick start

Nothing is downloaded until you ask for a sample, so looking costs seconds rather than 48 GB.

python
from datasets import load_dataset

covers = load_dataset("the-malware-files/pentimento-core", "covers", split="full", streaming=True)
print(next(iter(covers))["json"]["licence"])

stego = load_dataset("the-malware-files/pentimento-core", "wow-0050", split="full", streaming=True)

Each arm is its own config, named exactly as the arm table below names it, and covers is the default. Drop streaming=True to fetch a config to disk.

The single split is called full because the shards are not laid out along the train and test boundary. Cover records carry their own split field and SPLITS.md gives the rule; splitting any other way puts a cover in training and its own stego copy in test.

What makes it different

Three things, each of which exists because something else was missing.

Adaptive schemes alongside real tools. REVEAL, from the Netherlands Forensic Institute, covers more than 50 end-user hiding tools and excludes content-adaptive academic schemes on purpose. This corpus carries HUGO, WOW, S-UNIWARD, HILL, MiPOD, J-UNIWARD and UERD as separate labelled arms.

Per-source arms, not a blend. The ranking of embedding schemes is known to invert when the cover source changes. A corpus that blends sources hides exactly the effect a user most needs to measure, so the arms here stay separate and labelled.

Rebuildable byte for byte. Every generator is seeded and every file is recorded with a sha256 and the count of samples actually changed. Results in the steganalysis literature differ by several accuracy points on the choice of data split alone, same network and same algorithm, so a corpus that cannot be rebuilt identically cannot support a comparison.

Layout

Shards are WebDataset tar files. Inside each, a sample's parts share a basename:

pentimento-core-00000.tar
  000000.png     the image
  000000.json    its manifest row, licence included

They stream without unpacking, and every major dataset loader reads them.

PartFilesShardsWhat the rate means
Covers10,00010n/a
append_after_eoi-000010,00020not a rate: a fixed trailer
clean-grey10,00020n/a, clean
clean-jpeg10,00020n/a, clean
clean-jpeg-tools10,00020n/a, clean
clean-outguess8,11917n/a, clean
hill-005010,00020bits per pixel
hill-010010,00020bits per pixel
hill-020010,00020bits per pixel
hill-040010,00020bits per pixel
hugo-005010,00020bits per pixel
hugo-010010,00020bits per pixel
hugo-020010,00020bits per pixel
hugo-040010,00020bits per pixel
juniward-005010,00020bits per non-zero AC
juniward-010010,00020bits per non-zero AC
juniward-020010,00020bits per non-zero AC
juniward-040010,00020bits per non-zero AC
mipod-005010,00020bits per pixel
mipod-010010,00020bits per pixel
mipod-020010,00020bits per pixel
mipod-040010,00020bits per pixel
outguess-00508,11917fraction of capacity
outguess-02008,11917fraction of capacity
outguess-05008,11917fraction of capacity
steghide-005010,00020fraction of capacity
steghide-020010,00020fraction of capacity
steghide-050010,00020fraction of capacity
suniward-005010,00020bits per pixel
suniward-010010,00020bits per pixel
suniward-020010,00020bits per pixel
suniward-040010,00020bits per pixel
uerd-005010,00020bits per non-zero AC
uerd-010010,00020bits per non-zero AC
uerd-020010,00020bits per non-zero AC
uerd-040010,00020bits per non-zero AC
wow-005010,00020bits per pixel
wow-010010,00020bits per pixel
wow-020010,00020bits per pixel
wow-040010,00020bits per pixel

The four digits in an arm name are not one quantity. hugo-0200 is 0.2 bits per pixel; juniward-0200 is 0.2 bits per non-zero AC coefficient; steghide-0200 is 0.2 of the capacity steghide itself reports. Ranking arms on the number in the name compares three different things. Every sample's JSON record carries rate_unit, which is the authoritative answer per file.

Every JPEG arm is written at quality 95, and the adaptive spatial arms are simulated at the optimal embedding rate rather than by a real STC coder - the coding field on each record says so. Both are ordinary practice and both change what a result means, so neither should be discovered after the fact.

Each part ships its own checksum file, SHA256SUMS-covers and SHA256SUMS-arms, so verifying a download is one command:

sha256sum -c SHA256SUMS-covers

Do it before use: a shard that arrived truncated reads as a smaller corpus rather than as an error.

On Kaggle the shards are unpacked, and that command does not apply there. Kaggle extracts archives when they are uploaded and offers no way to refuse, so pentimento-core-00000.tar arrives as a folder of the same members under the same names. The bytes are the same; the container is gone. Verify that copy against each record's own checksum instead, which is a finer check because it names the file that is actually wrong:

python load_pentimento.py --verify pentimento-core-00000/

load_pentimento.py reads a folder and a tar the same way, so nothing else changes. The Internet Archive and HuggingFace copies are tar shards as described above.

Tiers nest. Nano is the first 200 covers of the same ordering Lite's first 1,000 and Core's 10,000 follow, so you can develop against a small tier and evaluate on a larger one without the two overlapping in a way that flatters the result.

Licensing, in one paragraph

Every file carries its own licence, and each file's own `.json` member inside the shards is the authoritative record of it. ATTRIBUTION.csv beside this file is an extract of those records for the covers that require a credit line, and is the one to read if you want the licences without downloading the shards. (This paragraph used to point at "the manifest", which does not ship as a standalone file; a reader who wanted to check a licence before committing to 48 GB was sent to something they could not find.) The collection is published as CC BY 4.0, which is the strictest obligation present, not the loosest: complying with it satisfies every file here. Where a file's own record names something looser, rely on that instead.

LicenceCovers
CC02,625
CC BY 2.02,624
Public domain1,922
CC BY 4.01,794
CC BY 3.0850
CC BY 2.5183
CC BY 1.02

5,453 of 10,000 covers (54.5%) require attribution, and each carries a ready-made credit line in its manifest row under attribution. Stego images are derivatives and inherit their cover's terms; every stego sample JSON carries the cover's licence under cover_licence, so a reader holding only one arm can still discharge the obligation.

See LICENCES.md for the full statement.

Before you train on it

Read SPLITS.md. A cover and its stego versions are far more alike than any two unrelated photographs, so a random split puts a cover in training and its own stego copy in test, and the classifier learns the photograph rather than the payload. Split by cover, never by image.

Citing

See CITATION.cff, which GitHub and Zenodo both render automatically.

Provenance

Covers come from Wikimedia Commons under permissive licences, deduplicated by perceptual hash across sources and sessions, and capped per photographer and per camera body so no single prolific uploader dominates the sample. Rejected candidates are recorded with their reason rather than silently dropped.

A caveat on that last sentence, kept because it was not true once. Cover SELECTION records its rejections. The arm BUILDERS did not: a failure during embedding was written to stderr and the loop continued, so 118 covers were missing from every spatial arm with nothing in the corpus saying so, and the build log that held the reason had been superseded. The images turned out to have been written without their manifest rows, they were rebuilt, and the arms are complete. The builders now record what they skip.