Team Ai
Apppublic

berencarkci/visual-analytics-assistant-v2

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes
App README

Multi-Agent Visual Analytics Assistant with Preference Optimization

Upload tabular data (CSV / Excel / JSON), ask a question in natural language, and get an appropriate chart plus an insight grounded in computed statistics.

Two modes

Single-agent — one model call produces the whole answer. This is the prompt-only baseline the project measures against. The model never sees the data, only the schema, so its insight is unverifiable and shown as raw JSON rather than as a finding.

Multi-agent — supervisor, data analyst, visualization and insight agents run in sequence, with a deterministic evaluation agent reviewing the result. The transform is executed on the real data, so insights quote numbers that were actually computed, and every number is verified mechanically before display. The Agent Trace tab shows each step.

Model

Qwen2.5-3B-Instruct with a LoRA adapter trained in two stages:

  • —SFT (QLoRA) on ~2,975 examples across five output formats (single-call, supervisor, data analyst, visualization, insight). Training data is generated by importing each agent's own prompt builder, so a training example is byte-identical to what the agent sends at inference.
  • —DPO (QLoRA) on 430 preference pairs was trained and evaluated, but did not improve held-out accuracy. That result is reported rather than hidden; the gains in this task came from targeted SFT data instead. The deployed adapter is the SFT model.

Evaluation

Measured on a frozen 60-question benchmark (38 dev / 22 test) plus a 25-question robustness probe for input shapes the benchmark does not cover — categorical and boolean columns, missing values, non-English column names, tiny tables. The system was also run against the independently collected NLV Corpus (real utterances from 102 participants) and completed 57 of 60 superstore utterances.

Notes

  • —Chart choices pass through deterministic guardrails that check the prepared data (category counts, negative values, discrete axes) and correct the model where a rule is violated. Corrections are shown, not hidden.
  • —The assistant does not forecast or explain causes. Asked to, it shows the relevant history rather than inventing an answer.
  • —Running on ZeroGPU: the first request after an idle period includes model load time. Multi-agent mode makes four model calls, so it is slower than single-agent — the trade-off buys verified, statistics-backed insights.