datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fashion-products-small-multimodal-embeddingsslay-the-spire-2-card-multimodal-embeddings
Slay the Spire 2: Multimodal Card Embeddings
Joint text+image embeddings for every card in Slay the Spire 2 (Early Access), produced by Qwen/Qwen3-VL-Embedding-2B. One unit-normalized 1024-D vector per card. Mechanically AND visually similar cards land near each other; cards across STS1 and STS2 share the coordinate system.
This is the multimodal-embeddings dataset. For text-only embeddings or the underlying card metadata + portraits, see:
t22000t/slay-the-spire-2-cards - metadata… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/slay-the-spire-2-card-multimodal-embeddings.slay-the-spire-1-card-multimodal-embeddings
Slay the Spire 1: Multimodal Card Embeddings
Joint text+image embeddings for every card in Slay the Spire (1.0 release), produced by Qwen/Qwen3-VL-Embedding-2B. One unit-normalized 1024-D vector per card. Mechanically AND visually similar cards land near each other; cards across STS1 and STS2 share the coordinate system.
This is the multimodal-embeddings dataset. For text-only embeddings or the underlying card metadata + portraits, see:
t22000t/slay-the-spire-1-cards - metadata +… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/slay-the-spire-1-card-multimodal-embeddings.t5-gemma-2-multimodal-embeddinggame-rep-bench-multimodal-embeddings
GameRepBench Multimodal Embeddings
Multi-view representations from title/description/metadata, icon, and multiple
store screenshots.
Rows: 1000
Text model: BAAI/bge-small-en-v1.5 (384D)
Visual model: Qdrant/clip-ViT-B-32-vision (512D)
Fused dimension: 896
Fusion: weighted normalized concatenation (0.55 text / 0.45 visual)
Maximum screenshots per app: 6
Rows with visual features: 1000
Apple rows with visual features: 999
Google rows with visual features: 1
Distinct visual views… See the full description on the dataset page: https://huggingface.co/datasets/kargaranamir/game-rep-bench-multimodal-embeddings.enriched-movie-dataset-with-multimodal-embeddings
Enriched Movie Dataset with Multimodal Embeddings
Dataset Description
This dataset provides rich metadata for over 44,000 movies, with a primary focus on providing a pre-computed, high-quality multimodal content embedding for each film.
It was created by fusing two popular Kaggle datasets: "The Movies Dataset" and the "IMDB Multimodal Vision & NLP Genre Classification" dataset. It has been further enriched with parsed text features and a unique 512-dimensional vector… See the full description on the dataset page: https://huggingface.co/datasets/ujwal-jibhkate/enriched-movie-dataset-with-multimodal-embeddings.
