Mitchins/anime-style-classifier-v5
Anime Style Classifier v5
An EfficientNet-B2 model that classifies anime frames into 4 production style categories and doubles as a style embedding model — capturing the visual drawing style of an image independently of scene content.
Model Details
Classes
Style Examples
<table> <tr> <td align="center"><b>flat</b></td> <td align="center"><b>modern</b></td> </tr> <tr> <td><img src="images/exampleflat0.jpg" width="400"/></td> <td><img src="images/examplemodern0.jpg" width="400"/></td> </tr> <tr> <td><img src="images/exampleflat1.jpg" width="400"/></td> <td><img src="images/examplemodern1.jpg" width="400"/></td> </tr> <tr> <td align="center"><b>painterly</b></td> <td align="center"><b>retro</b></td> </tr> <tr> <td><img src="images/examplepainterly0.jpg" width="400"/></td> <td><img src="images/exampleretro0.jpg" width="400"/></td> </tr> <tr> <td><img src="images/examplepainterly1.jpg" width="400"/></td> <td><img src="images/exampleretro1.jpg" width="400"/></td> </tr> </table>
Performance
OOD Evaluation (439 held-out images, never seen during training)
Overall accuracy: 98.41% (432/439)
Confusion Matrix
Error Analysis (7 total misclassifications)
All errors are plausible borderline cases. Click any thumbnail to view full size.
<table> <tr><th>Image</th><th>True</th><th>Predicted</th><th>Confidence</th><th>Notes</th></tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernpainterly18086660.jpg"><img src="images/errormodernpainterly18086660.jpg" width="120"/></a></td> <td>modern</td><td>painterly</td><td>95.8%</td> <td>Soft-lit scene with textured background</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernretro6883863.jpg"><img src="images/errormodernretro6883863.jpg" width="120"/></a></td> <td>modern</td><td>retro</td><td>98.5%</td> <td>Retro-styled modern production</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errormodernretro9338738.jpg"><img src="images/errormodernretro9338738.jpg" width="120"/></a></td> <td>modern</td><td>retro</td><td>93.3%</td> <td>Muted palette resembling cel-era</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernComfyUI00007.jpg"><img src="images/errorpainterlymodernComfyUI00007.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>96.0%</td> <td>AI-generated with clean digital finish</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernComfyUI00008.jpg"><img src="images/errorpainterlymodernComfyUI00008.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>79.5%</td> <td>Borderline painterly/modern</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorpainterlymodernscene959firstloverealizationpainterly00001.jpg"><img src="images/errorpainterlymodernscene959firstloverealizationpainterly00001.jpg" width="120"/></a></td> <td>painterly</td><td>modern</td><td>91.1%</td> <td>Synthetic painterly with modern elements</td> </tr> <tr> <td><a href="https://huggingface.co/Mitchins/anime-style-classifier-v5/resolve/main/images/errorretromodern16143039.jpg"><img src="images/errorretromodern16143039.jpg" width="120"/></a></td> <td>retro</td><td>modern</td><td>66.3%</td> <td>Low confidence — genuinely ambiguous era</td> </tr> </table>
Use as Style Embedding Model
Beyond classification, this model produces 1408-dimensional style embeddings from the penultimate layer that capture the visual drawing style of anime frames — line weight, shading technique, color palette, compositing approach — independently of scene content (characters, backgrounds, objects).
Embedding Performance
Evaluated on 1,800 frames from 18 anime series across all 4 style groups (100 frames per series):
6.5× better style separation than both ImageNet baselines. The fine-tuned model clusters same-style series together while pushing different styles apart. ImageNet models treat everything as nearly identical (0.82 cross-style similarity), unable to distinguish drawing style from scene content.
Embedding Usage
import torch
import timm
from PIL import Image
from torchvision import transforms
from safetensors.torch import load_file
# Load model
model = timm.create_model('efficientnet_b2', pretrained=False, num_classes=4)
state_dict = load_file('model.safetensors')
model.load_state_dict(state_dict)
model.eval()
transform = transforms.Compose([
transforms.Resize((512, 512)),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
def get_style_embedding(image_path):
"""Extract a 1408-dim style embedding from the penultimate layer."""
img = Image.open(image_path).convert('RGB')
x = transform(img).unsqueeze(0)
with torch.no_grad():
features = model.forward_features(x) # [1, 1408, 16, 16]
embedding = features.mean(dim=[-2, -1]) # [1, 1408] global avg pool
embedding = embedding / embedding.norm(dim=1, keepdim=True) # L2 normalize
return embedding.squeeze(0).numpy()
# Compare two frames
emb_a = get_style_embedding('frame_a.jpg')
emb_b = get_style_embedding('frame_b.jpg')
similarity = emb_a @ emb_b # cosine similarity (embeddings are L2-normalized)
print(f'Style similarity: {similarity:.4f}')
# > 0.8 = very similar style (likely same series/studio)
# 0.4-0.8 = same style family (e.g., both modern)
# < 0.2 = different style axis (e.g., retro vs modern)Interpreting Embedding Distances
The global average pooled features from forward_features() (before the classifier head) encode style properties:
The embedding captures how something is drawn — line weight, shading technique, color palette, compositing — not what is drawn. Scenes with completely different content (landscapes vs close-ups vs action sequences) from the same series will still cluster together.
Classification Usage
import torch
import timm
from PIL import Image
from torchvision import transforms
from safetensors.torch import load_file
# Load model
model = timm.create_model('efficientnet_b2', pretrained=False, num_classes=4)
state_dict = load_file('model.safetensors')
model.load_state_dict(state_dict)
model.eval()
# Preprocessing — resize to 512×512
transform = transforms.Compose([
transforms.Resize((512, 512)),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
image = Image.open('anime_frame.jpg').convert('RGB')
x = transform(image).unsqueeze(0)
with torch.no_grad():
logits = model(x)
probs = torch.softmax(logits, dim=1)
pred_idx = logits.argmax(-1).item()
confidence = probs[0, pred_idx].item()
classes = ['flat', 'modern', 'painterly', 'retro']
print(f'{classes[pred_idx]}: {confidence:.1%}')Training Details
Architecture and Training Pipeline
- Base training: EfficientNet-B2 pretrained on ImageNet, fine-tuned at 440px on style dataset with focal loss and class-balanced sampling
- Progressive resolution: Fine-tuned from 440→512px with frozen batch norm, discriminative learning rates (backbone 1/10th of head), and 1-epoch warmup
- Weight merging: Linear interpolation of base (440px) and progressive (512px) checkpoints at optimal α=0.54, lifting accuracy beyond either parent
Training Data
Split: 17,649 train / 1,962 validation (90/10) + 439 separate OOD test images.
Key Training Decisions
Focal loss handles class imbalance (retro has 4× more samples than flat). The best single fine-tuned checkpoint reached 96.36% OOD accuracy; linear weight merging of two complementary checkpoints pushed this to 98.41%. Batch norm freezing during progressive resolution training was critical — without it, running stat drift degraded OOD performance.
What This Model is NOT
- ❌ Not a character design classifier — won't detect moe, chibi, bishounen, etc.
- ❌ Not a content classifier — sees how it's drawn, not what
- ❌ Not an era detector — a modern show using cel technique → retro (correctly)
- ❌ Not an AI detector — AI art that nails a style gets classified by that style
- ❌ Not a quality scorer — good and bad modern frames both classify as modern
Limitations
- Trained primarily on Japanese anime (TV series, films). May not generalize to Western animation, donghua, or manhwa.
- "Painterly" class is hardest — the boundary with modern is subjective when digital tools simulate traditional media.
- AI-generated anime art (e.g., ComfyUI, Stable Diffusion) often confuses the model between painterly and modern.
- Flat style class has the fewest real-world examples — supplemented with curated synthetic images.
Image Disclaimer
Example and evaluation images shown in this model card are either generated via diffusion models or sourced from web search results for illustrative purposes only. Web-sourced images remain the property of their respective copyright holders and are used here solely as examples to demonstrate model behavior.
Citation
@misc{anime-style-classifier-v5,
title={Anime Style Classifier V5},
year={2026},
publisher={HuggingFace},
note={EfficientNet-B2 for anime production style classification and embedding}
}