Waheed786dar/Comiman-Dataset
Comiman Dataset Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset A license-gated comics and manga page corpus built for training a model that can plan and draw full comic/manga series (the planned model: Waheed786dar/Comiman). Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a public-domain claim limited to works… See the full description on the dataset page: https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset.
Comiman Dataset
Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset
A license-gated comics and manga page corpus built for training a model that can plan and draw full comic/manga series (the planned model: Waheed786dar/Comiman). Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a public-domain claim limited to works published up to 1963) and every page passed quality, NSFW, blank-page and duplicate filters. Every accept/reject decision is public in ledger/.
Dataset summary
By license type
By language (as tagged by the source)
By decade
By genre tag (keyword-derived, multi-label)
By quality tier
By reading direction
Curation decisions (ledger)
Supported tasks
Comic page / panel generation, layout and panel-order modeling, page-type and quality classification, retrieval with the included CLIP embeddings, panel detection (COCO json), and as the visual stage of a story-to-comic pipeline. Not included yet: OCR text, captions, speaker labels (planned GPU stage 2).
Dataset structure
data/{train,validation,test}-NNNNN.parquet page rows (images as WebP bytes + annotations)
webdataset/{train,validation,test}-NNNNN.tar same pages as WebDataset tar (key.webp + key.json)
coco/{split}-<tag>.json panel boxes in COCO format (absolute pixels)
manifest/books-<tag>.parquet one row per book
ledger/ledger-<tag>.parquet every accept/reject decision with reason
state/page_hashes-<tag>.parquet perceptual hashes (duplicate detection / resume)
integrity/files-<tag>.json sha256 + row counts of every shard
exports/books.jsonl, books.csv, ledger.csv convenience exportsPage fields
Book fields (manifest)
bookid, sourceid, split, title, creator, year, decade, language, subjects, genretags, readingdirection, licensetype, licenseconfidence, licenseurl, rightstext, copyrightstatus, attributionrequired, sharealike, attributiontext, source, sourceurl, fileused, pagecount, pagesdropped, storypages, webpbytes, avgpanels, meanquality, meanaesthetic, meansharpness, meannsfw, bookrating, ratingstars, qualitytier, added_at
Quality ratings (how they are computed)
Page quality (0-100) = weighted mean of: resolution of the source scan (25%), sharpness via Laplacian variance (25%), contrast (15%), ink coverage sanity (10%), panel structure found (10%), aesthetic head (15%, when available; weights are renormalized otherwise). Stars: >=85 five, >=70 four, >=55 three, >=40 two, else one.
Book rating (0-100) = 0.60 x mean page quality + 0.15 x completeness (kept / kept+dropped) + 0.10 x license confidence (high 100, medium 60) + 0.15 x share of pages with 2+ panels. Tier A >= 80, B >= 65, C >= 50, D otherwise. Books below 35.0 are not published.
How the data was built
- Discovery: Internet Archive search by comics/manga subjects plus license/copyright metadata; optional user-supplied archives listed in licenses.json.
- License gate: CC0/PDM/CC BY/CC BY-SA accepted; NC/ND rejected; "public domain" claims accepted only with a publication year <= 1963; everything else is rejected and logged.
- Extraction: CBZ, CBR (via libarchive) and PDF (embedded scan or rendered at 130 dpi), natural page order.
- Cleaning: junk/tiny images removed, scan-border auto-crop, resize to 1600 px, WebP q82, blank-page removal, perceptual-hash duplicate detection (book dropped when >= 0.8 of pages already exist).
- Annotation: panel boxes (OpenCV heuristic), image metrics, CLIP ViT-L/14 embeddings, aesthetic score, page-type, NSFW score on GPU (T4 x2).
- Validation before every upload: row counts, schema, image checksum + decode sample, tar member count, upload path whitelist; after upload remote file sizes are compared with local sizes.
Usage
from datasets import load_dataset
import io, numpy as np
from PIL import Image
ds = load_dataset("Waheed786dar/Comiman-Dataset", split="train", streaming=True)
row = next(iter(ds))
img = Image.open(io.BytesIO(row["image_webp"]))
emb = np.frombuffer(row["clip_l14_fp16"], dtype=np.float16) # (768,)
# only high quality story pages
good = ds.filter(lambda r: r["quality_star"] >= 4 and r["page_type"] in (None, "story"))WebDataset: load_dataset("webdataset", data_files="hf://datasets/Waheed786dar/Comiman-Dataset/webdataset/train-*.tar", split="train", streaming=True).
Book-level filtering: read manifest/ (or exports/books.csv), keep qualitytier in ("A","B"), then select pages by bookid.
Licensing and attribution
Compilation layer (metadata, ratings, boxes, embeddings, docs): CC BY 4.0 - attribution to Waheed is required (see LICENSE).
Underlying page images keep their original terms (license_type per book). Public-domain works: no copyright is claimed by the compiler.
CC BY books: keep attributiontext. CC BY-SA books (sharealike = true): share-alike applies.
Trained-model credit is requested ("Trained on the Comiman Dataset by Waheed"); strict enforcement against model weights is legally unsettled, so treat this as a condition of use for the data and a request for the model.
Considerations
License verification is metadata-level. Uploaders on source sites can be wrong. Public-domain status of old comics is a United States judgment (for example non-renewal) and may differ in your country. Report problems for takedown.
Historical content can contain outdated and offensive depictions. Review before training or serving.
Bias/coverage: strongly skewed to mid-century Western comics; little modern manga because modern manga is almost always copyrighted.
Heuristics: panel boxes, genre tags and reading direction are heuristic, not human-verified.
NSFW filter is a classifier and will miss some content and wrongly flag some art.
No personal data is collected; creators' names come from public archive metadata.
Citation
@misc{comiman_dataset,
title = {Comiman Dataset},
author = {Waheed},
year = {2026},
url = {https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset}
}Contact / takedown
GitHub issues: https://github.com/uzairlovesM/comiman-dataset-reports
