Anzhc/SDXL-Text-Encoder-Longer-CLIP-L
734
An experiment, with CLIP L trained with up to 770 tokens with ~10k anime dataset, without adjusting arch. Concatenation is used to accummulate features.


Token-adjusted, with images removed if they ca'nt meet token criteria:


