Team Ai
Datasetpublic

HuggingFaceM4/the_cauldron

Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. Load the dataset To load the dataset, install the library datasets with pip install datasets. Then, from datasets import load_dataset ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d") to download… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/the_cauldron.

sourceHugging Faceupdated 2y agoView on Hugging Face
560likes362kdownloads
../
filetrain-00000-of-00012-31f9e352dd92b89f.parquet472.2 MBdownload
filetrain-00001-of-00012-fd5209f92cc41a19.parquet459.6 MBdownload
filetrain-00002-of-00012-92679b2734ac7628.parquet504.6 MBdownload
filetrain-00003-of-00012-a8b56992c61d81ef.parquet465.6 MBdownload
filetrain-00004-of-00012-99a200654bed3175.parquet417.8 MBdownload
filetrain-00005-of-00012-6b1c18f2643c07ae.parquet496.2 MBdownload
filetrain-00006-of-00012-24c23a98e7115d64.parquet548.9 MBdownload
filetrain-00007-of-00012-a0ffeb3aaa88a42b.parquet563.1 MBdownload
filetrain-00008-of-00012-63a22f9c33ecebfa.parquet543.8 MBdownload
filetrain-00009-of-00012-28e0317bd329a609.parquet505.0 MBdownload
filetrain-00010-of-00012-08bd3953e91ca0a1.parquet429.4 MBdownload
filetrain-00011-of-00012-5110ca02b6f250c0.parquet483.3 MBdownload

HuggingFaceM4/the_cauldron · main · files are served by the source, never re-hosted here