Rithvik762/vln-trajectory-memory-stage2
VLN Trajectory-Memory — Stage 2 (projector alignment) Text-only question answering where the only source of truth is a robot's action history. Each record gives a navigation trajectory as a list of primitive actions and asks something that can only be answered by tracking where those actions lead: how far from the start, which way the robot faces, what happened in the last quarter of the route. It was built to measure whether a frozen vision-language model (Qwen3-VL-2B) can read… See the full description on the dataset page: https://huggingface.co/datasets/Rithvik762/vln-trajectory-memory-stage2.
VLN Trajectory-Memory — Stage 2 (projector alignment)
Text-only question answering where the only source of truth is a robot's action history. Each record gives a navigation trajectory as a list of primitive actions and asks something that can only be answered by tracking where those actions lead: how far from the start, which way the robot faces, what happened in the last quarter of the route.
It was built to measure whether a frozen vision-language model (Qwen3-VL-2B) can read a trajectory memory injected as a single token by a small trained projector, and to compare seven ways of injecting it. No images, no simulator — every answer is a deterministic function of action_codes (or, for the next category, of the step that followed).
Configs
Seven arms share the same rows, the same train/validation split and the same schema; they differ in how a trajectory is cut up before the encoder reads it. A row_index config lists the sampled rows and their split.
Columns
Composition
All seven arms cover the same 24,000 rows (train + validation), drawn to balance history length — which the source corpus does not — across RxR 16,167, Human 5,206, R2R 2,627.
Load
from datasets import load_dataset
ds = load_dataset("Rithvik762/vln-trajectory-memory-stage2", "continuous_final") # default config
val = load_dataset("Rithvik762/vln-trajectory-memory-stage2", "continuous_final", split="validation")
print(val[0]["question"], "->", val[0]["answer_key"])How it was built
tools/gen_projector_data.py samples rows from the NaVILA R2R / RxR / Human annotation files (which carry each step's executed action history), draws questions per category with a fixed seed, and splits by trajectory. Questions and answers are generated programmatically, never by a model. Full pipeline, training recipe and results: see the project's Stage-2 documentation.
Caveats
- Question text is not bit-reproducible. Draws use Python's
hash()of the scheme name, which is randomised per process unlessPYTHONHASHSEEDis fixed. Rows and the split are. - Balanced by length, not natural. Use
sample_weightto recover corpus proportions. - Float tolerance is strict. Distance answers are scored within 0.25 m, which is tight for RxR routes tens of metres long.
- Short histories agree across arms by construction. Under 60 actions every absolute arm and
continuous_finalproduce one identical read-out, which makes those rows a built-in control.
Provenance and licence
Derived from the NaVILA dataset mix (R2R, RxR and human-collected VLN-CE trajectories), which build on Matterport3D. The action histories and questions here are generated by this project, but the instruction text originates from those datasets — their terms govern redistribution. Check them before making any copy of this data public.
