Team Ai
Datasetpublic

Waheed786dar/Comiman-Dataset

Comiman Dataset Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset A license-gated comics and manga page corpus built for training a model that can plan and draw full comic/manga series (the planned model: Waheed786dar/Comiman). Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a public-domain claim limited to works… See the full description on the dataset page: https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset.

sourceHugging Faceotherupdated 3d agoView on Hugging Face
0likes889downloads
Dataset Card

Comiman Dataset

Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset

A license-gated comics and manga page corpus built for training a model that can plan and draw full comic/manga series (the planned model: Waheed786dar/Comiman). Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a public-domain claim limited to works published up to 1963) and every page passed quality, NSFW, blank-page and duplicate filters. Every accept/reject decision is public in ledger/.

Dataset summary

Books193
Pages50,467
Panels (heuristic)93,816
Image size on disk (WebP)5.58 GB
Mean book rating73.51 / 100
Splitstrain / validation / test (deterministic, by book, 96/2/2)
Last build2026-10-03

By license type

license_typebooks
publicdomainclaimed193

By language (as tagged by the source)

languagebooks
eng89
fre74
ger13
spa9
rus2
dut2
English1
por1
dan1
ita1

By decade

decadebooks
190041
189027
188026
191026
185024
187021
184010
18606
19205
18302
17901
17201
17001
19501
17801

By genre tag (keyword-derived, multi-label)

genrebooks
humor188
war15
unclassified4
children2
funny_animal1
romance1

By quality tier

tierbooks
B127
A43
C23

By reading direction

directionbooks
ltr193

Curation decisions (ledger)

decisionbooks
accepted193
rejected_year23
rejectedlowquality2
error_process1

Supported tasks

Comic page / panel generation, layout and panel-order modeling, page-type and quality classification, retrieval with the included CLIP embeddings, panel detection (COCO json), and as the visual stage of a story-to-comic pipeline. Not included yet: OCR text, captions, speaker labels (planned GPU stage 2).

Dataset structure

data/{train,validation,test}-NNNNN.parquet        page rows (images as WebP bytes + annotations)
webdataset/{train,validation,test}-NNNNN.tar      same pages as WebDataset tar (key.webp + key.json)
coco/{split}-<tag>.json                           panel boxes in COCO format (absolute pixels)
manifest/books-<tag>.parquet                      one row per book
ledger/ledger-<tag>.parquet                       every accept/reject decision with reason
state/page_hashes-<tag>.parquet                   perceptual hashes (duplicate detection / resume)
integrity/files-<tag>.json                        sha256 + row counts of every shard
exports/books.jsonl, books.csv, ledger.csv        convenience exports

Page fields

fieldtypemeaning
book_idstringstable id cmn-<sha1[:12]> of the source id
page_noint32reading page index inside the book (after filtering)
splitstringtrain / validation / test (by book, never split inside a book)
width, heightint32stored image size (longest side <= 1600)
origwidth, origheightint32size of the source scan before crop/resize
aspectratio, isspreadfloat32, boolwidth/height; spread = aspect > 1.15
croppedboolscan borders/margins auto-cropped
color_modestringcolor / grayscale / bw
reading_directionstringltr or rtl (manga, Arabic, Hebrew, Persian, Urdu)
image_webpbinaryWebP-encoded page
webp_sha256stringchecksum of image_webp
dhashuint6464-bit perceptual hash
panelcount, panelboxesint16, listheuristic panel boxes [x0,y0,x1,y1] normalized 0-1, reading order
sharpness, contrast, brightness, ink_ratiofloat32classical image metrics
aestheticfloat32LAION aesthetic head on CLIP ViT-L/14 (about 1-10; null if GPU model unavailable)
nsfw_scorefloat32probability from an NSFW classifier (pages above 0.85 were dropped)
pagetype, pagetype_confstring, float32zero-shot CLIP: cover / story / ad / textpage / blank / backcover
qualityscore, qualitystarfloat32, int8page rating 0-100 and 1-5 stars
clipl14fp16binaryL2-normalized CLIP ViT-L/14 embedding, 768 x float16

Book fields (manifest)

bookid, sourceid, split, title, creator, year, decade, language, subjects, genretags, readingdirection, licensetype, licenseconfidence, licenseurl, rightstext, copyrightstatus, attributionrequired, sharealike, attributiontext, source, sourceurl, fileused, pagecount, pagesdropped, storypages, webpbytes, avgpanels, meanquality, meanaesthetic, meansharpness, meannsfw, bookrating, ratingstars, qualitytier, added_at

Quality ratings (how they are computed)

Page quality (0-100) = weighted mean of: resolution of the source scan (25%), sharpness via Laplacian variance (25%), contrast (15%), ink coverage sanity (10%), panel structure found (10%), aesthetic head (15%, when available; weights are renormalized otherwise). Stars: >=85 five, >=70 four, >=55 three, >=40 two, else one.

Book rating (0-100) = 0.60 x mean page quality + 0.15 x completeness (kept / kept+dropped) + 0.10 x license confidence (high 100, medium 60) + 0.15 x share of pages with 2+ panels. Tier A >= 80, B >= 65, C >= 50, D otherwise. Books below 35.0 are not published.

How the data was built

  1. 1.Discovery: Internet Archive search by comics/manga subjects plus license/copyright metadata; optional user-supplied archives listed in licenses.json.
  2. 2.License gate: CC0/PDM/CC BY/CC BY-SA accepted; NC/ND rejected; "public domain" claims accepted only with a publication year <= 1963; everything else is rejected and logged.
  3. 3.Extraction: CBZ, CBR (via libarchive) and PDF (embedded scan or rendered at 130 dpi), natural page order.
  4. 4.Cleaning: junk/tiny images removed, scan-border auto-crop, resize to 1600 px, WebP q82, blank-page removal, perceptual-hash duplicate detection (book dropped when >= 0.8 of pages already exist).
  5. 5.Annotation: panel boxes (OpenCV heuristic), image metrics, CLIP ViT-L/14 embeddings, aesthetic score, page-type, NSFW score on GPU (T4 x2).
  6. 6.Validation before every upload: row counts, schema, image checksum + decode sample, tar member count, upload path whitelist; after upload remote file sizes are compared with local sizes.

Usage

python
from datasets import load_dataset
import io, numpy as np
from PIL import Image

ds = load_dataset("Waheed786dar/Comiman-Dataset", split="train", streaming=True)
row = next(iter(ds))
img = Image.open(io.BytesIO(row["image_webp"]))
emb = np.frombuffer(row["clip_l14_fp16"], dtype=np.float16)       # (768,)

# only high quality story pages
good = ds.filter(lambda r: r["quality_star"] >= 4 and r["page_type"] in (None, "story"))

WebDataset: load_dataset("webdataset", data_files="hf://datasets/Waheed786dar/Comiman-Dataset/webdataset/train-*.tar", split="train", streaming=True).

Book-level filtering: read manifest/ (or exports/books.csv), keep qualitytier in ("A","B"), then select pages by bookid.

Licensing and attribution

Compilation layer (metadata, ratings, boxes, embeddings, docs): CC BY 4.0 - attribution to Waheed is required (see LICENSE).

Underlying page images keep their original terms (license_type per book). Public-domain works: no copyright is claimed by the compiler.

CC BY books: keep attributiontext. CC BY-SA books (sharealike = true): share-alike applies.

Trained-model credit is requested ("Trained on the Comiman Dataset by Waheed"); strict enforcement against model weights is legally unsettled, so treat this as a condition of use for the data and a request for the model.

Considerations

License verification is metadata-level. Uploaders on source sites can be wrong. Public-domain status of old comics is a United States judgment (for example non-renewal) and may differ in your country. Report problems for takedown.

Historical content can contain outdated and offensive depictions. Review before training or serving.

Bias/coverage: strongly skewed to mid-century Western comics; little modern manga because modern manga is almost always copyrighted.

Heuristics: panel boxes, genre tags and reading direction are heuristic, not human-verified.

NSFW filter is a classifier and will miss some content and wrongly flag some art.

No personal data is collected; creators' names come from public archive metadata.

Citation

@misc{comiman_dataset,
  title  = {Comiman Dataset},
  author = {Waheed},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset}
}

Contact / takedown

GitHub issues: https://github.com/uzairlovesM/comiman-dataset-reports