Team Ai
Datasetpublic

Rithvik762/vln-trajectory-memory-stage2

VLN Trajectory-Memory — Stage 2 (projector alignment) Text-only question answering where the only source of truth is a robot's action history. Each record gives a navigation trajectory as a list of primitive actions and asks something that can only be answered by tracking where those actions lead: how far from the start, which way the robot faces, what happened in the last quarter of the route. It was built to measure whether a frozen vision-language model (Qwen3-VL-2B) can read… See the full description on the dataset page: https://huggingface.co/datasets/Rithvik762/vln-trajectory-memory-stage2.

sourceHugging Faceotherupdated 19d agoView on Hugging Face
0likes70downloads
Dataset Card

VLN Trajectory-Memory — Stage 2 (projector alignment)

Text-only question answering where the only source of truth is a robot's action history. Each record gives a navigation trajectory as a list of primitive actions and asks something that can only be answered by tracking where those actions lead: how far from the start, which way the robot faces, what happened in the last quarter of the route.

It was built to measure whether a frozen vision-language model (Qwen3-VL-2B) can read a trajectory memory injected as a single token by a small trained projector, and to compare seven ways of injecting it. No images, no simulator — every answer is a deterministic function of action_codes (or, for the next category, of the step that followed).

Configs

Seven arms share the same rows, the same train/validation split and the same schema; they differ in how a trajectory is cut up before the encoder reads it. A row_index config lists the sampled rows and their split.

configtrainvalidationhow the memory is read
continuous_final21,6022,398whole history, read once — the arm Stage 3 used
absolute_continuous_checkpoints21,6022,398one never-reset pass, read every 60 actions
absolute_reset_chunks21,6022,398independent 60-action chunks, state re-zeroed
absolute_hybrid_chunks21,6022,398chunks + dead-reckoned global pose
proportional_continuous_checkpoints21,6022,398same, boundaries at 25/50/75/100 %
proportional_reset_chunks21,6022,398same, proportional boundaries
proportional_hybrid_chunks21,6022,398same, proportional boundaries
row_index24,000—the sampled rows and their split

Columns

columnmeaning
action_codesthe trajectory: 0 stop, 1 forward 0.25 m, 2 turn left 15°, 3 turn right 15°
n_actions, length_binhistory length, and its bin (<=60, 61-120, 121-180, 181-220, >220)
question, answerthe prompt and the supervised target sentence
answer_key, key_kindwhat a scorer compares against, and how: int exact · float within 0.25 m · signed_deg within 15° · text exact
answer_valuethe same key as a number, so numeric answers can be compared without parsing; null when key_kind is text
categoryfacts (counts, first/last action) · spatial (heading, distance, displacement) · segment (a span of the route) · next (the action that followed)
instructionthe human navigation instruction — context only; no answer depends on it
dataset, trajectory_id, row_idprovenance; the split is by trajectory_id, so no trajectory appears in both
sample_weightnatural frequency of the row's length bin ÷ rows drawn from it (bins are balanced, so the corpus is not)

Composition

All seven arms cover the same 24,000 rows (train + validation), drawn to balance history length — which the source corpus does not — across RxR 16,167, Human 5,206, R2R 2,627.

categoryrows
spatial7,050
next6,166
segment5,978
facts4,806
length binrows
>2204,800
121-1804,800
181-2204,800
61-1204,800
<=604,800

Load

python
from datasets import load_dataset

ds = load_dataset("Rithvik762/vln-trajectory-memory-stage2", "continuous_final")          # default config
val = load_dataset("Rithvik762/vln-trajectory-memory-stage2", "continuous_final", split="validation")
print(val[0]["question"], "->", val[0]["answer_key"])

How it was built

tools/gen_projector_data.py samples rows from the NaVILA R2R / RxR / Human annotation files (which carry each step's executed action history), draws questions per category with a fixed seed, and splits by trajectory. Questions and answers are generated programmatically, never by a model. Full pipeline, training recipe and results: see the project's Stage-2 documentation.

Caveats

  • —Question text is not bit-reproducible. Draws use Python's hash() of the scheme name, which is randomised per process unless PYTHONHASHSEED is fixed. Rows and the split are.
  • —Balanced by length, not natural. Use sample_weight to recover corpus proportions.
  • —Float tolerance is strict. Distance answers are scored within 0.25 m, which is tight for RxR routes tens of metres long.
  • —Short histories agree across arms by construction. Under 60 actions every absolute arm and continuous_final produce one identical read-out, which makes those rows a built-in control.

Provenance and licence

Derived from the NaVILA dataset mix (R2R, RxR and human-collected VLN-CE trajectories), which build on Matterport3D. The action histories and questions here are generated by this project, but the instruction text originates from those datasets — their terms govern redistribution. Check them before making any copy of this data public.