aplominski/flickr30k-multilingual
Flickr 30k Multilingual Multilingual version of the Flickr30k dataset. This dataset contains images paired with original English captions, along with synthetic translations into Polish, German, French, and Spanish generated using Meta's NLLB-200-1.3B model. ๐ Quickstart You can easily load this dataset using the Hugging Face datasets library: from datasets import load_dataset # Load the dataset dataset = load_dataset("aplominski/flickr30k-multilingual") #โฆ See the full description on the dataset page: https://huggingface.co/datasets/aplominski/flickr30k-multilingual.
Flickr 30k Multilingual
Multilingual version of the Flickr30k dataset. This dataset contains images paired with original English captions, along with synthetic translations into Polish, German, French, and Spanish generated using Meta's NLLB-200-1.3B model.
๐ Quickstart
You can easily load this dataset using the Hugging Face datasets library:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("aplominski/flickr30k-multilingual")
# Inspect an example
sample = dataset["train"][0]
print("English Captions:", sample["captions_en"])
print("Polish Captions:", sample["captions_pl"])
# Display the image
sample["image"].show()๐ Dataset Structure
Data Fields
- `image`: A
PIL.Imageobject containing the image. - `captions_en`: A list of original English descriptions (5 per image).
- `captions_pl`: A list of Polish translated captions.
- `captions_de`: A list of German translated captions.
- `captions_es`: A list of Spanish translated captions.
- `captions_fr`: A list of French translated captions.
Data Splits
Splits are based on Karpathy split collected from (Delphboy/karpathy-splits github repostiory)[https://github.com/Delphboy/karpathy-splits].
๐ ๏ธ Translation Details
The translations were generated programmatically using the NLLB-200 (1.3B parameters) machine translation model based on the original English annotations.
๐ Citation & Credits
This dataset builds upon the original Flickr30k dataset:
Young, P., Lai, A., Hodosh, M., & Hockenmaier, J. (2014). From image descriptions to visual denotations: New richly annotated datasets for semantic training. Transactions of the Association for Computational Linguistics.
This dataset uses the standard Karpathy split based on:
Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3128โ3137.
โ๏ธ License
Distributed under the CC BY 4.0 license like orginal dataset.
