Team Ai
Datasetpublic

aplominski/flickr30k-multilingual

Flickr 30k Multilingual Multilingual version of the Flickr30k dataset. This dataset contains images paired with original English captions, along with synthetic translations into Polish, German, French, and Spanish generated using Meta's NLLB-200-1.3B model. ๐Ÿš€ Quickstart You can easily load this dataset using the Hugging Face datasets library: from datasets import load_dataset # Load the dataset dataset = load_dataset("aplominski/flickr30k-multilingual") #โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/aplominski/flickr30k-multilingual.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes115downloads
Dataset Card

Flickr 30k Multilingual

Multilingual version of the Flickr30k dataset. This dataset contains images paired with original English captions, along with synthetic translations into Polish, German, French, and Spanish generated using Meta's NLLB-200-1.3B model.

๐Ÿš€ Quickstart

You can easily load this dataset using the Hugging Face datasets library:

python
from datasets import load_dataset

# Load the dataset
dataset = load_dataset("aplominski/flickr30k-multilingual")

# Inspect an example
sample = dataset["train"][0]
print("English Captions:", sample["captions_en"])
print("Polish Captions:", sample["captions_pl"])

# Display the image
sample["image"].show()

๐Ÿ“Š Dataset Structure

Data Fields

  • โ€”`image`: A PIL.Image object containing the image.
  • โ€”`captions_en`: A list of original English descriptions (5 per image).
  • โ€”`captions_pl`: A list of Polish translated captions.
  • โ€”`captions_de`: A list of German translated captions.
  • โ€”`captions_es`: A list of Spanish translated captions.
  • โ€”`captions_fr`: A list of French translated captions.

Data Splits

SplitNumber of Examples
Train29,769
Validation1,014
Test1,000

Splits are based on Karpathy split collected from (Delphboy/karpathy-splits github repostiory)[https://github.com/Delphboy/karpathy-splits].

๐Ÿ› ๏ธ Translation Details

The translations were generated programmatically using the NLLB-200 (1.3B parameters) machine translation model based on the original English annotations.

๐Ÿ“ Citation & Credits

This dataset builds upon the original Flickr30k dataset:

Young, P., Lai, A., Hodosh, M., & Hockenmaier, J. (2014). From image descriptions to visual denotations: New richly annotated datasets for semantic training. Transactions of the Association for Computational Linguistics.

This dataset uses the standard Karpathy split based on:

Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3128โ€“3137.

โš–๏ธ License

Distributed under the CC BY 4.0 license like orginal dataset.