tojpaj/chart-qa-vlm-multilingual
Chart-QA Multilingual VLM (Hindi/Punjabi) — v2
LoRA adapter for google/gemma-3-4b-it-VLM, fine-tuned for native-language chart interpretation. AutoScientist Challenge Part 2, Data Visualization.
v2 result — +25.5% relative over v1
Trained via client.autoscientist.create().
How the config was chosen
v1 used AutoScientist's auto-selected hyperparameters (r=16). Inspecting the public adapter_config.json of four independent Challenge entrants — across Qwen3.5-0.8B, gpt-oss-20b, Llama-4-Scout-17B and Mistral-7B — showed every one had chosen `r=64`, alpha 128–256, dropout 0.05. Adopting that config lifted the win rate from 0.5599 to 0.6818 on the first iteration; four further search iterations added only +0.021 combined.
The static config delivered ~87% of the gain; search depth delivered ~13%.
best_hyperparams returned the pinned LoRA values unchanged, confirming that explicit hyperparams survive AutoScientist's search rather than being overridden — the search tuned learning rate and scheduler around them.
No overfitting at 3 epochs
Another entrant documented r=64, 3 epochs on 20k rows of medical reasoning producing a model that lost to its base model (58/42), with eval-loss plateauing after epoch 1. That did not reproduce here:
eval_loss 1.6267 → 1.5103 → 1.4469 → 1.4066 → 1.3986 (monotonically down)
train_loss 11.40 → 1.48Different task (chart VQA vs. open-ended reasoning) and scale (1,260 vs 20k rows). Competitor configs are a prior worth testing, not a rule — the per-iteration curve decides.
Data
1,260 rows (420 English + 420 Hindi + 420 Punjabi) over 100 synthetic charts (bar/line/grouped-bar/stacked-bar/scatter) with deterministic ground-truth answers — see tojpaj/chart-qa-multilingual-indic.
Limitations
- Numeric answers are visual estimates (e.g. 77.0 vs a ground truth 76.7) — expected for chart reading, not a defect.
- Charts are synthetic and clean; performance on real dashboards or infographics is untested.
- Win rate is AutoScientist's internal metric, not an independent benchmark.
