datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polyu-storyworld-charactersMixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.character_select_stand_alone_apphttps://github.com/mirabarukaso/character_select_stand_alone_app
game_characters
Database of Characters in Mobile Games
All the character in the following games are supported:
Arknights (crawled from https://prts.wiki)
Fate/Grand Order (crawled from https://fgo.wiki)
Azur Lane (crawled from https://wiki.biligame.com/blhx)
Girls' Front-Line (crawled from https://iopwiki.com/)
Genshin Impact (crawled from https://genshin-impact.fandom.com/ja/wiki/%E5%8E%9F%E7%A5%9E_Wiki)
The source code and python library is hosted on narugo1992/gchar, and the scheduled job is… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_characters.Kohaku-Delta-Alpha-Characters-Result
Test Index for Kohaku-Delta-Alpha Model
ID
Tag
Copyright
Gender
Posts
CCIP
AIC
BP
Core Tags
2075918
inkling_player_character
splatoon_(series)
female
9646
0.383333
0.998157
0.645165
inkling girl, long hair, tentacle hair, pointy ears, bangs, blunt bangs, red eyes
2040387
doodle_sensei_(blue_archive)
blue_archive
female
6632
0.721212
0.992072
0.625653
halo, bangs, blue eyes, breasts, long hair, black hair, blue hair, hair ornament
1978860
2b_(nier:automata)… See the full description on the dataset page: https://huggingface.co/datasets/AngelBottomless/Kohaku-Delta-Alpha-Characters-Result.mixamo-characters
Mixamo Characters Dataset
This dataset contains 3D character models from Mixamo in FBX format.
Contents
Format: FBX (Autodesk Filmbox)
Pose: T-Pose
Rig: Mixamo auto-rig (compatible with Mixamo animations)
Total characters: 108
File Structure
<character_name>.fbx
Usage
from huggingface_hub import hf_hub_download
# Download a specific character
filepath = hf_hub_download(
repo_id="your-username/mixamo-characters",
filename="Abe.fbx"… See the full description on the dataset page: https://huggingface.co/datasets/GbotHQ/mixamo-characters.Vivarium-Greenscreen-Characters-FLUX-9B
Vivarium Greenscreen Characters — FLUX 9B
This dataset contains fictional, illustrated foreground characters made for animation, games, storyboards and other creative projects. People are shown full-length against a solid green background so they can be separated from a scene and composited over a location image. The visual style is warm, hand-painted and gently anime-inspired; it is not photorealistic.
What you will find
The collection includes 500 people with… See the full description on the dataset page: https://huggingface.co/datasets/laion/Vivarium-Greenscreen-Characters-FLUX-9B.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.fictional-characters-image-datasetHow to use
Here is how to use this dataset:
from datasets import load_dataset
dataset = load_dataset("gryffindor-ISWS/fictional-characters-image-dataset")
This repository contains fictional characters dataset constructed from Wikidata for the research project "Draw Me Like Your Triples: Leveraging Generative AI for the Completion of Wikidata". The project was conducted by Raia Abu Ahmad, Martin Critelli, Şefika Efeoğlu, Eleonora Mancini, Célian Ringwald and Xinyue Zhang under the… See the full description on the dataset page: https://huggingface.co/datasets/gryffindor-ISWS/fictional-characters-image-dataset.generic_charactersqwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.character_similarity
character_similarity
This is a dataset used for training models to determine whether two anime images (containing only one person) depict the same character. The dataset includes the following versions:
Version
Filename
Characters
Images
Information
v0
images_v0.tar.xz
2059
162116
Crawled from zerochan.net, includes images of Arknights, Fate/Grand Order, Genshin Impact, Girls' Frontline, and Azur Lane, as well as over 1500 other game or anime characters. The images are… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_similarity.modular_characters_largemodular_charactersv2Roleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters.
The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.font_square_charactersDevanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.mahabharata-characters
Mahabharata Multidimensional Characters Dataset
An analytical, multidimensional dataset classifying 307 figures from the Sanskrit epic Mahabharata, combining genealogical and martial attributes with TypeSafe AI System One decision primitives (Choice, Score, Noul) and deterministic ethical indices.
Dataset Structure
data/train.jsonl: Full 307 character records in JSON Lines format.
data/characters.csv: Flattened tabular format with all ethical, martial, and… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/mahabharata-characters.ai-characters-QA
データ作成方法
以下の質問集のデータセットを利用しました。
質問に対する回答をLLMによって生成し、QA集のデータセットを新たに作成しました。各種キャラクターは、回答生成時のシステムプロンプトでキャラ付けしています。
回答生成には、llama-cpp-pythonとunsloth/gemma-3-27b-it-UD-Q8_K_XL.ggufを使用しました。
aituber_question_dataset
modular_characterscharacters_descriptionsmetrixel-rigged-characters
Metrixel Animated Character Samples
A small set of rigged, animated humanoid character models, free to use in your own projects. Each file ships with a skeleton, skin weights and a motion clip, so it drops straight into a scene — or retarget your own motion onto the rig instead.
They are published as complimentary sample content for Metrixel, a local 3D dataset-preparation toolchain — handy starting assets if you want to try a multi-view render, SDF or mesh-tensor export… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-rigged-characters.Charactersbooru-characters
Booru Characters
Overview
A line-oriented JSON dataset of character tag metadata extracted from Danbooru using the Danbooru API. Each record contains tag-level metadata and simple relationships between tags.
Contents
hf_dataset/characters.jsonl: one JSON object per line. Each object contains the fields described below.
hf_dataset/dataset_info.json: minimal metadata describing the exported features.
Fields (per record)
id (int):… See the full description on the dataset page: https://huggingface.co/datasets/Sn0w123/booru-characters.Taiwanese-Chinese_characters-POJ-Collectionannotations_creators:
expert-generated
language:
zh
en
language_creators:
expert-generated
license:
mit
multi-linguality:
monolingual
pretty_name: '
Taiwanese text dataset: a Chinese characters and POJ collection'
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-classification
feature-extraction
task_ids:
multi-label-classification
multi-class-classification
Gemini-3.1-Pro-GLM5-CharactersPrompts generated by Gemini 3.1 Pro.
Responses generated by Gemini 3.1 Pro.
Reasoning traces:
Step 1: Generated by GLM 5 which was provided the original system prompt / knowledge
Step 2: Edited by Gemini to fix any contradictions with the existing response
Step 3: Edited again to remove / reduce drafting and remove reasoning related to safety / refusals
System prompt was generated based on the original and whatever constraints / rules etc. had been mentioned in the reasoning trace.
A small… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-GLM5-Characters.SD_Anime_Characters_Repositorychub_popular_characters
