Elliot-Data/doclaynet_train_cleaned
doclaynet_train_cleaned The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 54,199 QA turns 145,370 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 43 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/doclaynet_train_cleaned.
037
No card is published for this repository, or it could not be fetched from Hugging Face right now.
