datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ComicsPAP
Comics: Pick-A-Panel
Updated val and test on 25/02/2025
This is the dataset for the ICDAR 2025 Competition on Comics Understanding in the Era of Foundational Models.
Please, check out our 🚀 arxiv paper 🚀 for more information 😊
The competition is hosted in the Robust Reading Competition website and the leaderboard is available here.
The dataset contains five subtask or skills:
Sequence Filling
Given a sequence of comic panels, a missing panel, and a set of option panels, the… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/ComicsPAP.ComicsPAP
Comics: Pick-A-Panel
Updated val and test on 25/02/2025
This is the dataset for the ICDAR 2025 Competition on Comics Understanding in the Era of Foundational Models.
Please, check out our 🚀 arxiv paper 🚀 for more information 😊
The competition is hosted in the Robust Reading Competition website and the leaderboard is available here.
The dataset contains five subtask or skills:
Sequence Filling
Given a sequence of comic panels, a missing panel, and a set of option panels, the… See the full description on the dataset page: https://huggingface.co/datasets/ansabgillani/ComicsPAP.comics-tools
FLUX workflows — c0sm1c_m1a (Mia)
Generation workflows for the FLUX.1-dev LoRA line. SDXL-era workflows remain in
workflows/ root (historical).
mia-flux-v1-test.json — LoRA smoke test
Minimal single-sampler graph, stock nodes only (no custom packs needed).
UI-format JSON: drag-drop onto the ComfyUI canvas or Import.
Graph: CheckpointLoaderSimple → LoraLoader → CLIPTextEncode (pos) +
CLIPTextEncode (neg, inert) + EmptyLatentImage → KSampler → VAEDecode → SaveImage… See the full description on the dataset page: https://huggingface.co/datasets/crimsonmythos/comics-tools.comics-audio-video
Comics Audio Video Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Audio Video data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/purdueinformatics/comics-audio-video.Star_Marvel_comics
Dataset Card for Star Villain Marvel Comics LoRa
Data set for Duskfallcrew/Star_Marvel_comics_LoRa
Trained with: https://colab.research.google.com/github/Linaqruf/kohya-trainer/blob/main/kohya-LoRA-dreambooth.ipynb
Where else can i find this ?
Both safetensors files are at: https://civitai.com/models/14831/star-ryan-ripley
The outputs aren't comic format.
I haven't tested it in WEB UI yet, the scripted outputs largely rely on the actual model.… See the full description on the dataset page: https://huggingface.co/datasets/EarthnDusk/Star_Marvel_comics.comics-audio-video-benchmark
Comics Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/josantos6/comics-audio-video-benchmark.comics_dataset_lineart_1024ComicsPonyXLdl-comics
Comics Image Depth Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Image Depth inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/watanabekelvin/dl-comics.comics-collection
Comics Image Depth Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Image Depth metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/wwojcikantoni/comics-collection.comicstrips-gpt4o-blip3
Comic Strips
Dataset Details
Dataset Description
This dataset contains indie comics from Reddit, then captioned with GPT4o and BLIP3.
Currently, only the GPT4o captions are available in this repository. The BLIP3 captions will be uploaded soon.
Roughly 1400 images were captioned at a cost of ~$11 using GPT4o (25 May 2024 version).
Curated by: @pseudoterminalx
Funded by @pseudoterminalx
License: MIT
Dataset Sources
Unlike other free-to-use… See the full description on the dataset page: https://huggingface.co/datasets/bghira/comicstrips-gpt4o-blip3.nlp-comics
Comics Image Audio Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Image Audio data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/rraoswati/nlp-comics.paper-comics-2023
Comics Pointcloud Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Pointcloud Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/yazeedalha/paper-comics-2023.comics-sensor-fusion-mini
Comics Sensor Fusion Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Sensor Fusion inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/amritastatistics04/comics-sensor-fusion-mini.postdoc-comics-2024
Comics Text Tabular Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Text Tabular metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/LevPopov/postdoc-comics-2024.comics_dataset_512_inv_manga_correctedcomic-setupcomics-pointcloud-text-v2-2024
Comics Pointcloud Text Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Pointcloud Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/madisonjones0910/comics-pointcloud-text-v2-2024.homework-comics
Comics Video Text Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Video Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and usage… See the full description on the dataset page: https://huggingface.co/datasets/mwilliamsdale/homework-comics.comics-collection
Comics Audio Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Audio Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/brunoymartins/comics-collection.comics-video-text
Comics Video Text Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Video Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/joshuathomaswood/comics-video-text.comics-audio-video-clean
Comics Audio Video Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Audio Video data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/edwardswilliam/comics-audio-video-clean.comics-image-depth
Comics Image Depth Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Image Depth metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/briangonzalezora/comics-image-depth.comics-image-text
Comics Image Text Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Image Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/filipgrabowski/comics-image-text.comics-sensor-fusion
Comics Sensor Fusion Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Sensor Fusion data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/jonas-neumann/comics-sensor-fusion.comics-audio-video-benchmark
Comics Audio Video Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Audio Video metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/darrenhuasaw/comics-audio-video-benchmark.comics_orig_lineartcomics-corpus
Comics Image Audio Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Image Audio metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/Felixschmidt/comics-corpus.comics-corpus
Comics Multimodal3 Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Multimodal3 data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/Arjunchopra/comics-corpus.comics-image-text
Comics Image Text Data Notes
Dataset summary
Preparation notes and schema examples for Comics tasks using Image Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/ggarniersebastien1987/comics-image-text.
