opencompass/TextEdit
TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models Danni Yang, Sitao Chen, Changyao Tian If you find our work helpful, please give us a โญ or cite our paper. See the InternVL-U technical report appendix for more details. ๐ News [2026/03/06] TextEdit benchmark released. [2026/03/06] Evaluation code and initial baselines released. [2026/03/06] Leaderboard updated with latest models. ๐โฆ See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.
93.9k
1#!/bin/bash2path="your_project_path_here"3cd "$path/TextEdit"4 5python eval_pipeline/vlm_metrics_eval_step2.py \6 --answer_dir "$path/TextEdit/result/vlm_gemini_full_answers" \7 --output_file "$path/TextEdit/result/gemini_report_fullset.json" \8 --weights 0.4 0.3 0.1 0.1 0.1 \9 --enable_cutoff10 11python eval_pipeline/vlm_metrics_eval_step2.py \12 --answer_dir "$path/TextEdit/result/vlm_gemini_mini_answers" \13 --output_file "$path/TextEdit/result/gemini_report_miniset.json" \14 --weights 0.4 0.3 0.1 0.1 0.1 \15 --enable_cutoff16 