Team Ai
Modelpublic

Mitchins/anime-style-classifier-v5

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes67downloads
Model Card

Anime Style Classifier v5

An EfficientNet-B2 model that classifies anime frames into 4 production style categories and doubles as a style embedding model — capturing the visual drawing style of an image independently of scene content.

Model Details

ArchitectureEfficientNet-B2 (timm)
Parameters7.7M
Input512 × 512 RGB
Embedding dim1408
Classes4 (flat, modern, painterly, retro)
FormatSafeTensors
File size30 MB
Training data19,611 images across 4 classes
Training methodSupervised CE + progressive resolution fine-tuning (440→512) + weight merge

Classes

StyleDescriptionExample series
flatMinimal shading, solid color blocks, low gradientsPing Pong The Animation, The Heike Story, Kaiba
modernHigh-detail digital pipeline, clean gradients, contemporary AAADemon Slayer, Violet Evergarden, Attack on Titan
painterlyVisible brush texture, watercolor/oil feel, textured transitionsMushi-Shi, Your Lie in April, Land of the Lustrous
retroCel-era aesthetic, grain, simpler compositing, limited dynamic rangeCowboy Bebop, Neon Genesis Evangelion, Robotech

Style Examples

<table> <tr> <td align="center"><b>flat</b></td> <td align="center"><b>modern</b></td> </tr> <tr> <td><img src="images/exampleflat0.jpg" width="400"/></td> <td><img src="images/examplemodern0.jpg" width="400"/></td> </tr> <tr> <td><img src="images/exampleflat1.jpg" width="400"/></td> <td><img src="images/examplemodern1.jpg" width="400"/></td> </tr> <tr> <td align="center"><b>painterly</b></td> <td align="center"><b>retro</b></td> </tr> <tr> <td><img src="images/examplepainterly0.jpg" width="400"/></td> <td><img src="images/exampleretro0.jpg" width="400"/></td> </tr> <tr> <td><img src="images/examplepainterly1.jpg" width="400"/></td> <td><img src="images/exampleretro1.jpg" width="400"/></td> </tr> </table>

Performance

OOD Evaluation (439 held-out images, never seen during training)

Overall accuracy: 98.41% (432/439)

ClassRecallCountErrors
flat100.0%83/830
modern98.5%203/2063
painterly94.0%47/503
retro99.0%99/1001

Confusion Matrix

→ flat→ modern→ painterly→ retro
flat83000
modern020312
painterly03470
retro01099

Error Analysis (7 total misclassifications)

All errors are plausible borderline cases. Click any thumbnail to view full size.

<table> <tr><th>Image</th><th>True</th><th>Predicted</th><th>Confidence</th><th>Notes</th></tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernpainterly18086660.jpg"><img src="images/errormodernpainterly18086660.jpg" width="120"/></a></td> <td>modern</td><td>painterly</td><td>95.8%</td> <td>Soft-lit scene with textured background</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernretro6883863.jpg"><img src="images/errormodernretro6883863.jpg" width="120"/></a></td> <td>modern</td><td>retro</td><td>98.5%</td> <td>Retro-styled modern production</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernretro9338738.jpg"><img src="images/errormodernretro9338738.jpg" width="120"/></a></td> <td>modern</td><td>retro</td><td>93.3%</td> <td>Muted palette resembling cel-era</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernComfyUI00007.jpg"><img src="images/errorpainterlymodernComfyUI00007.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>96.0%</td> <td>AI-generated with clean digital finish</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernComfyUI00008.jpg"><img src="images/errorpainterlymodernComfyUI00008.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>79.5%</td> <td>Borderline painterly/modern</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernscene959firstloverealizationpainterly00001.jpg"><img src="images/errorpainterlymodernscene959firstloverealizationpainterly00001.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>91.1%</td> <td>Synthetic painterly with modern elements</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorretromodern16143039.jpg"><img src="images/errorretromodern16143039.jpg" width="120"/></a></td> <td>retro</td><td>modern</td><td>66.3%</td> <td>Low confidence — genuinely ambiguous era</td> </tr> </table>

Use as Style Embedding Model

Beyond classification, this model produces 1408-dimensional style embeddings from the penultimate layer that capture the visual drawing style of anime frames — line weight, shading technique, color palette, compositing approach — independently of scene content (characters, backgrounds, objects).

Embedding Performance

Evaluated on 1,800 frames from 18 anime series across all 4 style groups (100 frames per series):

ModelIntra-Series Cohesion ↑Same-Style Similarity ↑Cross-Style Similarity ↓**Separation Gap** ↑
anime-style-classifier-v5 (this model)0.3470.8000.4640.336
Base EfficientNet-B2 (ImageNet)0.2280.8770.8260.051
ViT-B/16 (ImageNet)0.2140.8440.7950.049

6.5× better style separation than both ImageNet baselines. The fine-tuned model clusters same-style series together while pushing different styles apart. ImageNet models treat everything as nearly identical (0.82 cross-style similarity), unable to distinguish drawing style from scene content.

Style GroupCentroid CohesionNotes
flat0.835Yuasa / Science SARU / rotoscoped — tight cluster
modern0.906Digital AAA productions — very tight
painterly0.885Watercolor / textured — cohesive
retro0.624More varied (1985–2004 span) — intentional

Embedding Usage

python
import torch
import timm
from PIL import Image
from torchvision import transforms
from safetensors.torch import load_file

# Load model
model = timm.create_model('efficientnet_b2', pretrained=False, num_classes=4)
state_dict = load_file('model.safetensors')
model.load_state_dict(state_dict)
model.eval()

transform = transforms.Compose([
    transforms.Resize((512, 512)),
    transforms.ToTensor(),
    transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])

def get_style_embedding(image_path):
    """Extract a 1408-dim style embedding from the penultimate layer."""
    img = Image.open(image_path).convert('RGB')
    x = transform(img).unsqueeze(0)
    with torch.no_grad():
        features = model.forward_features(x)          # [1, 1408, 16, 16]
        embedding = features.mean(dim=[-2, -1])        # [1, 1408] global avg pool
        embedding = embedding / embedding.norm(dim=1, keepdim=True)  # L2 normalize
    return embedding.squeeze(0).numpy()

# Compare two frames
emb_a = get_style_embedding('frame_a.jpg')
emb_b = get_style_embedding('frame_b.jpg')
similarity = emb_a @ emb_b  # cosine similarity (embeddings are L2-normalized)
print(f'Style similarity: {similarity:.4f}')
# > 0.8  = very similar style (likely same series/studio)
# 0.4-0.8 = same style family (e.g., both modern)
# < 0.2  = different style axis (e.g., retro vs modern)

Interpreting Embedding Distances

The global average pooled features from forward_features() (before the classifier head) encode style properties:

Cosine SimilarityInterpretation
> 0.8Near-identical style (same series, same studio)
0.5 – 0.8Same style family (e.g., two modern AAA shows)
0.0 – 0.5Different style families
< 0.0Opposite ends of the style spectrum (e.g., 1980s cel vs 2020s digital)

The embedding captures how something is drawn — line weight, shading technique, color palette, compositing — not what is drawn. Scenes with completely different content (landscapes vs close-ups vs action sequences) from the same series will still cluster together.

Classification Usage

python
import torch
import timm
from PIL import Image
from torchvision import transforms
from safetensors.torch import load_file

# Load model
model = timm.create_model('efficientnet_b2', pretrained=False, num_classes=4)
state_dict = load_file('model.safetensors')
model.load_state_dict(state_dict)
model.eval()

# Preprocessing — resize to 512×512
transform = transforms.Compose([
    transforms.Resize((512, 512)),
    transforms.ToTensor(),
    transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])

image = Image.open('anime_frame.jpg').convert('RGB')
x = transform(image).unsqueeze(0)

with torch.no_grad():
    logits = model(x)
    probs = torch.softmax(logits, dim=1)
    pred_idx = logits.argmax(-1).item()
    confidence = probs[0, pred_idx].item()

classes = ['flat', 'modern', 'painterly', 'retro']
print(f'{classes[pred_idx]}: {confidence:.1%}')

Training Details

Architecture and Training Pipeline

  1. 1.Base training: EfficientNet-B2 pretrained on ImageNet, fine-tuned at 440px on style dataset with focal loss and class-balanced sampling
  2. 2.Progressive resolution: Fine-tuned from 440→512px with frozen batch norm, discriminative learning rates (backbone 1/10th of head), and 1-epoch warmup
  3. 3.Weight merging: Linear interpolation of base (440px) and progressive (512px) checkpoints at optimal α=0.54, lifting accuracy beyond either parent

Training Data

ClassImagesDescription
flat2,138Real anime frames + synthetic consensus images
modern3,954Extracted from 50+ modern anime series
painterly5,130Real frames + Safebooru/DeviantArt + validated synthetic
retro8,389Extracted from 30+ pre-digital-era series
Total19,611

Split: 17,649 train / 1,962 validation (90/10) + 439 separate OOD test images.

Key Training Decisions

Focal loss handles class imbalance (retro has 4× more samples than flat). The best single fine-tuned checkpoint reached 96.36% OOD accuracy; linear weight merging of two complementary checkpoints pushed this to 98.41%. Batch norm freezing during progressive resolution training was critical — without it, running stat drift degraded OOD performance.

What This Model is NOT

  • —❌ Not a character design classifier — won't detect moe, chibi, bishounen, etc.
  • —❌ Not a content classifier — sees how it's drawn, not what
  • —❌ Not an era detector — a modern show using cel technique → retro (correctly)
  • —❌ Not an AI detector — AI art that nails a style gets classified by that style
  • —❌ Not a quality scorer — good and bad modern frames both classify as modern

Limitations

  • —Trained primarily on Japanese anime (TV series, films). May not generalize to Western animation, donghua, or manhwa.
  • —"Painterly" class is hardest — the boundary with modern is subjective when digital tools simulate traditional media.
  • —AI-generated anime art (e.g., ComfyUI, Stable Diffusion) often confuses the model between painterly and modern.
  • —Flat style class has the fewest real-world examples — supplemented with curated synthetic images.

Image Disclaimer

Example and evaluation images shown in this model card are either generated via diffusion models or sourced from web search results for illustrative purposes only. Web-sourced images remain the property of their respective copyright holders and are used here solely as examples to demonstrate model behavior.

Citation

bibtex
@misc{anime-style-classifier-v5,
  title={Anime Style Classifier V5},
  year={2026},
  publisher={HuggingFace},
  note={EfficientNet-B2 for anime production style classification and embedding}
}