datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UniMER_Dataset
UniMER Dataset
For detailed instructions on using the dataset, please refer to the project homepage: UniMERNet Homepage
Introduction
The UniMER dataset is a specialized collection curated to advance the field of Mathematical Expression Recognition (MER). It encompasses the comprehensive UniMER-1M training set, featuring over one million instances that represent a diverse and intricate range of mathematical expressions, coupled with the UniMER Test Set, meticulously… See the full description on the dataset page: https://huggingface.co/datasets/wanderkid/UniMER_Dataset.UniMER
UniMER Dataset
For detailed instructions on using the dataset, please refer to the project homepage: UniMERNet Homepage
Introduction
The UniMER dataset is a specialized collection curated to advance the field of Mathematical Expression Recognition (MER). It encompasses the comprehensive UniMER-1M training set, featuring over one million instances that represent a diverse and intricate range of mathematical expressions, coupled with the UniMER Test Set, meticulously… See the full description on the dataset page: https://huggingface.co/datasets/deepcopy/UniMER.unimer-1mtest2_UniMERNetunimer_train_cleaned
unimer_train_cleaned
The unimer_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
542,449
QA turns
1,070,876
answers rewritten by the cleaning pass
138,530
QA created by the cleaning pass (new_qa)
528,261 (49.3%)
shards
3
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/unimer_train_cleaned.unimer_train_cleaned
unimer_train_cleaned
The unimer_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
542,449
QA turns
1,070,876
answers rewritten by the cleaning pass
138,530
QA created by the cleaning pass (new_qa)
528,261 (49.3%)
shards
3
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/unimer_train_cleaned.
