maticmatusek/VLM_semantics_SLO_benchmark
VLM Semantics SLO Benchmark VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia. The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.
VLM Semantics SLO Benchmark
VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia.
The released JSON contains 4,950 image-level records. Every record has ten Slovenian questions, giving 49,500 image-question task slots. Images are represented by source URLs; image files are not bundled in this release.
Dataset summary
Question types
Record structure
Each JSON object represents one image.
Example, shortened for readability:
{
"image_id": "1",
"image_url": "https://upload.wikimedia.org/.../image.JPG",
"image_categories": ["geografija", "naselja", "umetnost"],
"SLO": "Kaj nam ta slika pove o Sloveniji?",
"AnswerSLO": "...",
"F1": "Opiši glavni motiv slike in elemente, ki ga podpirajo v ozadju.",
"AnswerF1": "...",
"F1_suitable": "da"
}Data construction
Candidate images and metadata were collected from cultural collections and public web sources, followed by manual inspection, merging, deduplication, filtering, and metadata normalization. Images were assigned one or more categories from a Slovenian cultural taxonomy.
For each image, one Slovenian question was selected for each of the ten reasoning types. Gemini 2.5 Flash was used during benchmark construction to produce the reference answers and auxiliary suitability/correlation labels. These answers should therefore be treated as machine-generated silver-standard references, not as independently verified human ground truth.
The image_url field preserves image provenance at the URL level, but the release does not provide complete per-image licence metadata.
Loading the data
The JSON file can always be read with the Python standard library:
import json
with open("VLM_semantics_SLO_benchmark.json", encoding="utf-8") as file:
records = json.load(file)
print(len(records)) # 4950
print(records[0]["SLO"])The data use a wide format. To iterate over individual image-question pairs:
QUESTION_TYPES = ["SLO", "F1", "F2", "F3", "I1", "I2", "I3", "A1", "A2", "A3"]
for record in records:
for question_type in QUESTION_TYPES:
example = {
"image_id": record["image_id"],
"image_url": record["image_url"],
"question_type": question_type,
"question": record[question_type],
"answer": record[f"Answer{question_type}"],
}OCM Categories
Top category counts:
Intended uses
Suitable uses include:
- evaluation of Slovenian-capable VLMs;
- research on cultural grounding and computational semiotics;
- analysis of multimodal hallucination and visual specificity;
- supervised fine-tuning after appropriate filtering and licence review;
- comparison of literal and higher-level interpretive visual reasoning.
The dataset should not be treated as an authoritative source of cultural facts or used without expert review in high-stakes educational, heritage, legal, or social applications.
Limitations and biases
- Reference answers and metadata were generated partly by an LLM and may contain factual, visual, linguistic, or cultural errors.
- Cultural interpretations can be subjective and context-dependent.
- Web-source availability and source representation are uneven.
- Remote image links can produce link rot and make exact future reproduction difficult.
- The dataset is Slovenian-focused and does not represent every Slovenian region, community, identity, historical perspective, or contested interpretation.
- Public web images can contain identifiable people or sensitive contexts; users should review their intended use.
Licensing and image rights
The metadata and generated text are released for research use under the terms selected by the dataset publisher. The source images are not covered by a single dataset-wide licence. Copyright, attribution requirements, and permitted uses remain governed by each original source.
Before downloading, redistributing, training on, or publishing an image, verify its source page, licence, terms of use, and any applicable privacy or personality rights.
