Alibaba-DAMO-Academy/RynnValue-8B-Quantile
<h1 align="center"><a href="https://github.com/alibaba-damo-academy/RynnValue">RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance</a></h1>
<p align="center"> <a href="https://alibaba-damo-academy.github.io/RynnValue.github.io/">๐ Homepage</a> | <a href="https://github.com/alibaba-damo-academy/RynnValue">๐ป GitHub</a> | <a href="https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnvalue">๐ค HuggingFace</a> | <a href="https://www.modelscope.cn/collections/DAMO_Academy/RynnValue">๐ฎ ModelScope</a> | <a href="arxiv.org/abs/2608.09853">๐ ArXiv</a> </p>
A general-purpose value foundation model for robot manipulation โ and the full toolchain to evaluate it and use it for reinforcement learning of VLA policies.
RynnValue is a RynnBrain-based vision-language model (implemented on the Qwen3-VL architecture) that watches a robot video together with a task instruction and predicts, for every frame, the temporal distance โ the remaining time (in seconds) until the task is completed โ alongside a natural-language analysis of the trajectory (video description, instructionโvideo match, task success). Because temporal-distance labels are derived directly from timestamps, RynnValue scales to 7,000+ hours of heterogeneous embodied data (~3M instruction-conditioned clips) without any preference or progress annotations. The predicted time-to-completion is a dense, task-grounded value signal that can be used directly as a progress estimator, a reward model for policy evaluation and ranking, a critic for reinforcement learning of vision-language-action (VLA) policies, or an offline data curator that filters demonstration frames before behavior cloning.
Trained without a single preference label, RynnValue-8B reaches an average Kendall's ฯโ of 0.704 on RBM-EVAL-OOD โ above the fully preference-supervised state of the art (0.655) and more than double a progress-only counterpart (0.292).
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/RynnValue_arch.png" width="100%" alt="RynnValue overview"/> </p>
Overview. Given a language instruction and a sequence of sampled observations, RynnValue builds an interleaved multimodal sequence of repeated absolute-value (<value>) and relative-value (<relative_value>) query groups, encoded by the RynnBrain backbone in a single forward pass. Two distributional heads predict the absolute temporal distance to task completion and the signed relative temporal displacement between observations, while the LM head produces video analysis and task verification. The resulting temporal values serve as a unified interface for progress estimation, failure detection, and reward specification in robotic RL.
This repository bundles the complete stack:
Table of Contents
- Highlights
- How RynnValue Works
- Training Recipe
- Results
- Repository Structure
- Installation
- Quickstart: Inference
- Using RynnValue Programmatically
- Evaluation with Robometer
- Reward Server
- Policy RL with pi-rl
- Checkpoint Conversion
- Model Zoo
- Acknowledgements
- License
Highlights
- Temporal distance as the scaling target. Instead of preferences or normalized
[0,1]progress, RynnValue predicts the goal-conditioned cost-to-go in physical seconds. Labels come directly from timestamps (plus subtask segmentation and cutoff relabeling), so supervision scales to 7,000+ hours / ~3M clips across 10 heterogeneous data sources (AgiBot, EgoDex, Open X-Embodiment, RoboMIND, RoboTwin, โฆ) without a single preference or progress annotation. - State-of-the-art without preference labels. RynnValue-8B attains an average Kendall's ฯโ of 0.704 on RBM-EVAL-OOD (4B: 0.679), surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. That 4B is already this strong indicates the gains come from the temporal-distance formulation and its training/architectural designs rather than sheer model scale.
- Shortcut suppression by design. Random temporal sampling and temporal-order shuffling break the correspondence between sequence position / sampling interval and task progress; value-isolation attention (
pred_slot_isolated_eager) keeps each value-query group visible only to its own languageโvisual context, so predictions can't extrapolate from other value tokens. Ablations (average ฯโ, full model 0.704): removing shuffling โ 0.159, uniform sampling โ 0.207, removing value isolation at training time โ 0.611, removing the auxiliary language supervision โ 0.622, removing relative temporal-distance supervision โ 0.578. - Train isolated, infer causal. Value-isolation attention is a training-time constraint. At inference RynnValue is queried with a standard causal mask instead, which reuses the default attention kernels and โ on paired trajectories โ actually ranks better: 0.646 โ 0.704 at 8B and 0.614 โ 0.679 at 4B. Earlier value states appear to disambiguate the final queried score by carrying trajectory context. See Policy Ranking: Attention Mask and Value Tokenizer.
- Quantile value tokenizer. Frame-level remaining-time labels are extremely skewed (median 11.6 s, 90th percentile 68.7 s, max 636.3 s), so the absolute and relative heads use 256-bin quantile supports โ knots placed at the inverse empirical CDF of 2.17B training-frame labels, covering
[0, 636.3]s and (mirrored around zero)[โ636.3, 636.3]s. 119 of the 256 knots fall within 10 s and 217 within 50 s; fixed-interval centers spanning the same range put only ~20 of 256 within 50 s. Targets are two-hot encoded and decoded as expected bin centers; the support is frozen across training, fine-tuning, and evaluation, since each knot index is a fixed output-bin identity. - Language-grounded analysis. The model also generates an
Analysisblock: a video description, a Match: Yes/No verdict (does the video match the instruction?), and a Success: Yes/No verdict. Instruction-mismatch augmentation (10% of training samples) teaches the model to detect instructionโvideo mismatches instead of always reporting smooth progress. - A practical reward interface for real robots. Converted into dense rewards via potential-based shaping (
ฮฆ_t = โv_t), RynnValue raises real-world dual-arm Franka policy success from 52.5% โ 72.5% online and 63.8% โ 82.5% offline over the strongest reward-model baseline. - Value-based data curation. Used as an offline filter over pooled multi-task demonstrations, RynnValue lifts multi-task behavior-cloning success from 35.0% โ 42.5% with no target-task adaptation and no critic โ see Filter-BC.
- Full evaluation + RL loop. The bundled Robometer fork benchmarks RynnValue against 10+ reward-model baselines (RBM, GVL, ReWiND, RFM, RL-VLM-F, Robo-Dopamine, RoboReward, TopReward, VLAC, โฆ); the pi-rl fork closes the loop by improving ฯโ.โ policies with value-driven offline IQL and online SAC, plus a Filter-BC data-curation baseline that trains only on the frames RynnValue scores as making progress.
How RynnValue Works
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Instruction โโโโบโ โ
Metadata โโโโโโบโ RynnBrain backbone (RynnValueLangModel) โโโโบ Analysis text
Frames โโโโโโโโบโ โ (description / Match / Success)
โ hidden states at repeated query-token positions โ
โโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโ
โ <value> group (รN) โ <relative_value> group (รN)
โผ โผ
absolute value head relative value head
(256 quantile bins, two-hot) (256 quantile bins, two-hot)
โ โ
โผ โผ
remaining time per frame signed ฮt between adjacent framesThe architecture follows the paper (ยง3.2, Model Architecture). For each clip the input sequence is
x = [ m, โ, Iโ, Vโ, Iโ, Rโ, Vโ, โฆ, I_K, R_{Kโ1}, V_K, p_ver ]with embodiment metadata m, instruction + system prompts โ, K = 8 sampled observations I, an absolute-value query group Vแตข of N = 8 repeated <value> tokens per observation, a relative-value query group Rแตข of N = 8 repeated <relative_value> tokens between consecutive observations, and a trailing verification prompt p_ver.
- Grouped temporal queries. A single query token is an information bottleneck; each temporal prediction instead uses a group of N = 8 repeated query tokens that attend bidirectionally within the group, and whose hidden states are concatenated (not averaged) before the head, preserving complementary visual cues (object configuration, robotโobject interaction, task stage, completion evidence).
- Dual distributional heads. The absolute head predicts the remaining time to the (relabeled) completion cutoff; the relative head predicts the signed temporal displacement between consecutively presented observations โ adjacency in the multimodal sequence, not in the original video, so the target can be negative. Both heads are BroNet residual MLPs (hidden width 4096, depth 8, ReLU) mapping the concatenated
Nยทdgroup representation to 256 bin logits. Targets are two-hot encoded over frozen quantile supports and trained with cross-entropy (so gradients don't scale with temporal error); inference decodes the expected bin center in seconds. - Value-isolation attention. Query groups belonging to different observations cannot attend to each other, and later observation tokens cannot attend to earlier query tokens โ each temporal estimate must be grounded in the instruction and visual evidence, and value prediction never leaks into later groups. This mask is a training-time constraint: at inference RynnValue is queried with a standard causal mask, which needs no custom kernels and scores better on trajectory ranking (see Highlights).
- Language analysis & verification. A verification prompt after the last query group triggers autoregressive generation of
Video Description โ Match โ Successthrough the LM head. Success prediction is learned purely through this language objective โ there is no separate success-classification head.
The prompt (built by rynn_value/conversations.py) interleaves an optional embodiment/camera meta block, the task instruction, the two value questions, per-frame images followed by their <relative_value> / <value> token slots, and finally the analysis request. The processor (rynn_value/processing_rynn_value_lang.py) exposes a single-call inference API, process_episode(instruction, images, robot_description, camera_description).
Key implementation files:
The package self-registers with HuggingFace Auto classes (AutoConfig / AutoModel / AutoProcessor under model type rynn_value_lang), so exported checkpoints load with trust_remote_code=True and nothing else.
Training Recipe
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/RynnValue-train.png" width="100%" alt="RynnValue training pipeline and value-isolation attention"/> </p>
(a) Training strategy. Random temporal sampling and temporal-order shuffling suppress shortcuts tied to sampling intervals and sequence position, while instruction-mismatch augmentation strengthens languageโvisual grounding. RynnValue jointly learns absolute temporal distance, relative temporal displacement, and natural-language supervision. (b) Value-isolation attention. Within each value-query group, repeated queries attend to one another and to the languageโvisual context, while remaining isolated from other value-query groups. Colored cells denote visible attention connections.
Data. RynnValue is trained on a heterogeneous mixture of real-world, simulated, and egocentric trajectories โ 1.67M original episodes expanded into 3.09M instruction-conditioned segments (7,000+ hours, 223,395 instructions) via subtask segmentation and cutoff relabeling:
Long demonstrations are split using native temporal annotations where available (subtask); otherwise the whole episode is kept as a single coarse segment. Because recorded trajectories often contain post-completion motion, each segment is also assigned a completion cutoff โ the segment endpoint by default, with dataset-specific ratio- or duration-based trimming where that better approximates the first semantically completed observation.
Temporal-distance labels are then generated directly from timestamps: for frame t, cutoff c, and frame rate f, the target is v* = max((c โ t)/f, 0) seconds โ remaining time before the cutoff, zero at or after it. No dataset-specific progress normalization is needed. Qwen3-VL-27B captions supervise the Video Description output and do not affect the temporal targets.
Objectives. Three jointly optimized cross-entropy losses, L = L_abs + L_rel + ฮป_lang ยท L_lang with ฮป_lang = 2: (1) the absolute temporal-distance loss over two-hot bin targets, masked for instruction-mismatched samples (the original cutoff is no longer valid under a substituted instruction); (2) the relative temporal-distance loss, which is instruction-independent and therefore kept for mismatched samples; (3) the causal LM loss over the Video Description / Match / Success tokens.
Shortcut suppression. For each clip, K = 8 observations are sampled at irregular timestamps (random temporal sampling). Temporal-order shuffling is applied with probability 0.5: half of the sequences are independently sampled without sorting, the rest follow a forward-biased temporal walk with rewind probability 0.3 โ so relative targets can be negative. Together with value-isolation attention, this forces every prediction to be grounded in the corresponding observation and task semantics. For 10% of samples the instruction is swapped with one from a different trajectory (instruction-mismatch augmentation) and supervised toward Match: No / Success: No.
Optimization. AdamW with lr 1e-6, ฮฒ = (0.9, 0.95), weight decay 0.1, ฮต = 1e-8; constant LR schedule without warm-up, gradient norm clipped at 100. Training runs in bfloat16 with FSDP hybrid sharding and a per-device batch size of 2.
Quantile support. The 256 temporal knots are fit once from the frame-level labels of the full training mixture (2,167,034,629 labels, frame-weighted so longer trajectories contribute proportionally more) by sampling the empirical inverse CDF at equally spaced probability levels, with no upper-tail clipping โ the last knot is the maximum observed remaining time, 636.33 s. The signed relative support mirrors the same non-negative knots around zero. Knots are stored in float32 and nudged to the next representable value wherever consecutive quantile levels collide, keeping the support strictly increasing; it then stays frozen across training stages, fine-tuning runs, and checkpoint evaluation.
Reward interface. At inference, frames are fed chronologically and decoded into remaining time v_t. The potential ฮฆ_t = โv_t (negative before completion, approaching zero at the goal) preserves the physical temporal scale rather than normalizing to a task-specific [0,1] interval. Downstream RL uses it through potential-based shaping with a retained sparse completion term,
r'_t = ฮบ ยท (ฮณยทฮฆ_{t+1} โ ฮฆ_t) + { 0 if t = T and the trajectory succeeds, else โ1 }where ฮบ controls shaping strength (0.1 for RynnValue, 1.0 for Robometer in our real-world runs). The sparse term is kept because reward-model predictions can be noisy while the success label is clean.
Results
Policy ranking (RBM-EVAL-OOD). The suite contains 976 trajectories from six out-of-distribution datasets spanning different institutions, embodiments, camera viewpoints, and task families, each annotated with failed / suboptimal / successful execution quality levels. The ranking score is the absolute head's remaining time at the final queried observation, negated (โv_end), with instruction-mismatch gating: when the language branch predicts Match: No the score is replaced by โ100. Neither the relative head, the video description, nor the success verdict enters the score. Kendall's ฯโ is computed per task, averaged over tasks within a suite, then averaged unweighted over the six suites.
Trained without preference labels, RynnValue-8B reaches 0.704 and even RynnValue-4B reaches 0.679 โ both above the 0.655 of the full Robometer model trained with progress and trajectory-preference supervision, and far above its progress-only ablation (0.292). Among preference-free methods, RynnValue variants are best on all six suites, lifting the strongest prior preference-free average (0.502) to 0.679 / 0.704.
Kendall's ฯโ (โ). Bold marks the best overall result per column; baseline results are taken from Robometer.
Ablations. All variants are evaluated at the same training checkpoint; each row removes exactly one component from the full model.
Removing temporal-order shuffling is by far the most damaging (0.159, and negative on MIT Franka) โ sequence position becomes a strong proxy for progress again. Uniform sampling is nearly as bad (0.207): regular intervals let the model replay a stereotypical value curve instead of reading the visual content. The remaining three components each contribute 0.08โ0.13.
Instructionโtrajectory alignment. Scoring every instruction against every trajectory, RynnValue produces the clearest diagonal structure with the highest normalized diagonal margin (0.79 vs. 0.67 for the strongest baseline). All models are re-evaluated under a unified protocol from their publicly released weights:
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/Matrix.png" width="100%" alt="Instruction-trajectory confusion matrices"/> </p>
Value-curve quality. On real-world trajectories, RynnValue reacts sharply to task regressions and recoveries where normalized-progress baselines stay flat. Both curves are oriented so that higher means closer to completion (Robometer emits normalized progress, RynnValue the potential ฮฆ_t = โv_t, so the scales are not directly comparable):
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/value-case.png" width="100%" alt="Temporal-value curve comparison on a real-world trajectory"/> </p>
Three differences stand out: (1) in the highlighted regression interval RynnValue's potential drops sharply while Robometer responds weakly โ a consequence of order shuffling and value isolation grounding predictions in visual evidence; (2) after recovery RynnValue climbs steadily, whereas Robometer shows long plateaus and abrupt jumps โ the benefit of jointly learning absolute and relative temporal values; (3) RynnValue stays sensitive near completion, so late disturbances cause an immediate potential decrease followed by recovery instead of premature reward saturation.
Scaling behavior. Task diversity โ not episode volume โ drives generalization. Two subset families are trained from scratch under identical settings: one subsamples episodes within a fixed task set, the other subsamples tasks at fixed per-task episode counts, so both protocols see a comparable number of episodes at each fraction. Scaling episode count saturates almost immediately, while adding tasks monotonically reduces temporal-distance error on held-out unseen tasks well past the midpoint:
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/scaling_analysis.png" width="55%" alt="Scaling episode volume vs. task diversity"/> </p>
Real-world policy learning. Used as a zero-shot reward annotator (none of the tasks, objects, or scenes appear in training, and no target-domain fine-tuning) on a dual-arm Franka with wrist + two third-person cameras, across four manipulation tasks and 20 trials per task. Offline RL is IQL on mixed-expertise datasets (โ100 successful trajectories per task plus every unsuccessful attempt); online RL is DSRL latent steering, initialized from the per-task SFT checkpoint on Bread / Steak and from the Robometer offline-RL checkpoint on Box-in-Drawer / Bimanual, with that initialization held fixed across reward variants. Both use the same potential-based shaping interface with a retained sparse completion term.
Success rate (%). RynnValue is best on all four tasks in both regimes, with the largest margins on Steak Serving (+30 online) and Bimanual Box Transfer (+35 online). It is also more efficient: offline Bread Basket reaches 100% in 16.8 action chunks on average vs. 18.9 for Robometer, and Steak Serving / Box-in-Drawer both reach 90% in 14.9 chunks. RL with RynnValue solves Bimanual Box Transfer, on which multi-task SFT records no successes at all. Online gains on Box-in-Drawer are limited for both reward models (65.0 / 70.0, from the 50.0 Robometer offline checkpoint that both online runs start from): the task hinges on grasp stability and boxโdrawer alignment that third-person RGB alone does not disambiguate.
<p align="center"> <img src="https://raw.githubusercontent.com/alibaba-damo-academy/RynnValue/refs/heads/main/assets/figures/case_study.png" width="100%" alt="Representative real-world manipulation tasks"/> </p>
Scalable post-training via data filtering. Offline and online RL both need task-specific training, so RynnValue is also evaluated as a task-amortized data curator: demonstrations from all four tasks are pooled, every segment is scored, low-quality ones are dropped, and a single multi-task policy is fine-tuned on what survives. Same initialization, same source dataset, same optimization recipe as multi-task SFT โ the only difference is the filtering, which isolates the effect of value-based data selection.
Success rate (%). Filtering is not uniformly positive: Box-in-Drawer drops to zero, most likely the same visual-ambiguity failure as above โ with only third-person RGB, useful segments get scored as non-progressing and discarded. Mechanism and launcher in Filter-BC.
Repository Structure
RynnValue001/
โโโ rynn_value/ # RynnValue model package (HF trust_remote_code style)
โโโ rynn_infer/
โ โโโ inference.py # CLI: video + instruction โ value curve + analysis
โ โโโ plot_utils.py # save_video_with_trend(): renders annotated output video
โ โโโ outputs/ # sample inference outputs
โโโ robometer/ # Robometer benchmark fork (reward-model training + eval)
โ โโโ robometer/
โ โ โโโ configs/ # Hydra configs (reward_model/rynnvalue.yaml, distributed/fsdp.yaml, โฆ)
โ โ โโโ models/ # RBM (progress/preference/success heads), ReWiND transformer
โ โ โโโ evals/ # run_baseline_eval.py, eval/baseline servers, baselines/rynnvalue.py
โ โ โโโ data/ trainers/ # dataset + training code
โ โโโ rynnvalue_eval/ # RynnValue-specific eval launchers (policy ranking, confusion matrix, server)
โ โโโ eval_commands/ # eval commands for the other baselines
โ โโโ dataset_upload/ # converters to the RBM HF dataset format (LIBERO, AgiBotWorld, custom)
โ โโโ train.py # reward-model training entry (accelerate + FSDP, LoRA)
โโโ pi-rl/ # openpi fork + RL
โ โโโ src/openpi/ # ฯโ / ฯโ.โ
models (JAX + PyTorch), training, RL data loading
โ โโโ scripts/ # train.py, train_iql.py, serve_policy.py, launch scripts
โ โโโ configs/ # RL configs (SAC, TD, DAgger, EXPO)
โ โโโ examples/ # libero, droid, aloha, dsrl_sim, dsrl_franka, franka (real robot), โฆ
โ โโโ packages/openpi-client/ # websocket policy client
โ โโโ third_party/jaxrl2/ # vendored jaxrl2 (pi_iql, pixel_iql, pixel_sac)
โโโ tools/
โ โโโ convert_rynn_value_lang_to_hf.py # training ckpt โ standalone HF model
โโโ example/
โโโ Put_the_box_in_the_drawer_and_close_it.mp4Installation
Each component manages its own environment. Python 3.10 is required throughout.
RynnValue inference (rynn_value + rynn_infer)
Managed with uv via the top-level pyproject.toml:
uv sync # creates .venv with torch / transformers / imageio / โฆA recent transformers with Qwen3-VL support is required (pinned in pyproject.toml). A single GPU with โฅ 24 GB memory comfortably runs the 8B model in bf16.
Robometer evaluation
cd robometer
uv sync # Python 3.10, uv-managed (see pyproject.toml / uv.lock)See robometer/README.md and robometer/FINETUNE_ROBOMETER.md for dataset download, reward-model fine-tuning (LoRA + FSDP), and the full baseline matrix.
pi-rl
cd pi-rl
GIT_LFS_SKIP_SMUDGE=1 uv sync # JAX 0.5.3 (cuda12), flax, torch 2.7.1, lerobot, โฆGPU requirements follow upstream openpi: > 8 GB for inference, > 22.5 GB for LoRA fine-tuning, > 70 GB (A100/H100) for full fine-tuning. See pi-rl/README.md.
Quickstart: Inference
Run RynnValue on the bundled example video:
cd rynn_infer
uv run python inference.py \
--model_path Alibaba-DAMO-Academy/RynnValue-8B \
--video_path ../example/Put_the_box_in_the_drawer_and_close_it.mp4 \
--instruction "Put the box in the drawer and close it" \
--robot_description "a Franka single-arm robot" \
--camera_description "the main right camera" \
--num_frames 8 \
--output_path ./outputsUseful flags:
The script produces, in a timestamped output directory:
output_with_trend.mp4โ the input video with a synchronized Remaining Time (s) curve;- the parsed Analysis (video description,
Match: Yes/No,Success: Yes/No) and per-frame values.
Using RynnValue Programmatically
import torch
from transformers import AutoConfig, AutoModel, AutoProcessor
model_path = "/path/to/RynnValue-8B"
config = AutoConfig.from_pretrained(model_path, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_path, config=config, torch_dtype=torch.bfloat16,
trust_remote_code=True, device_map="cuda",
# Causal masking is what the reported numbers use. The value-isolation mask
# (`pred_slot_isolated_eager`) is a training-time device and ranks worse at
# inference; omitting this argument takes whatever `config.json` records.
attn_implementation="sdpa",
).eval()
processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
# `images`: list of PIL.Image frames sampled from the trajectory video
inputs = processor.process_episode(
instruction="Put the box in the drawer and close it",
images=images,
).to(model.device)
with torch.no_grad():
out = model(**inputs)
remaining_time = out.value.pred_value.float().mean(dim=0) # (num_frames,) seconds, averaged over value heads
delta_time = out.relative.pred_value # (num_frames,) signed time deltas
entropy = out.value.entropy.mean(dim=0) # (num_frames,) distributional uncertaintyEvaluation with Robometer
The robometer/ fork adds RynnValue as a first-class baseline (robometer/robometer/evals/baselines/rynnvalue.py, Hydra config reward_model=rynnvalue). Ready-made example commands live in robometer/rynnvalue_eval/ (plus start_server.sh for the reward server).
Policy ranking (does the value model rank better policies higher?) โ example commands in rynnvalue_eval/policy_ranking.sh; the released-model run:
cd robometer
ROBOMETER_PROCESSED_DATASETS_PATH=/path/to/processed_datasets \
python robometer/evals/run_baseline_eval.py \
'reward_model=rynnvalue' \
'model_path=Alibaba-DAMO-Academy/RynnValue-8B' \
'custom_eval.eval_types=[policy_ranking]' \
'custom_eval.policy_ranking=[rbm-1m-ood]' \
'custom_eval.use_frame_steps=false' \
'custom_eval.pad_frames=false' \
'custom_eval.num_examples_per_quality_pr=1000' \
'max_frames=8' \
'model_config.checkpoint_path=null' \
'model_config.mode=absolute' \
'model_config.stride=2' \
'model_config.num_frames=8' \
'model_config.attn_implementation=sdpa' \
'model_config.camera_desc_lookup_path=./extracted_meta_with_descriptions.json'Two optional model_config fields: camera_desc_lookup_path=extracted_meta_with_descriptions.json supplies the per-trajectory camera descriptions used by RBM-EVAL-OOD, and attn_implementation (sdpa / eager / pred_slot_isolated_eager) selects the inference-time attention mask. Set it explicitly rather than relying on the default: left unset, the run takes whatever the checkpoint's own config.json records, which is not the same across checkpoint generations. sdpa is the causal mask the reported numbers use; pred_slot_isolated_eager reproduces the value-isolation rows of the table below. To evaluate a fine-tuned raw checkpoint instead, set model_config.checkpoint_path=/path/to/checkpoint_model_XXXXXX (a model.pt directory whose sibling huggingface/ holds the processor/config artifacts).
Confusion matrix (cross-task instruction/video matching) โ example commands in rynnvalue_eval/confusion_matrix.sh, one per scoring mode (model_config.confusion_score_mode):
match_binaryโ score 1.0 iff the Analysis verdict isMatch: Yes;normalized_valueโ the value head normalized to[0, 1](1 โ t/t_max) on matches, 0 otherwise.
cd robometer
ROBOMETER_PROCESSED_DATASETS_PATH=/path/to/processed_datasets \
python robometer/evals/run_baseline_eval.py \
'reward_model=rynnvalue' \
'model_path=Alibaba-DAMO-Academy/RynnValue-8B' \
'custom_eval.eval_types=[confusion_matrix]' \
'custom_eval.confusion_matrix=[[...dataset ids...]]' \
'max_frames=8' \
'model_config.stride=2' \
'model_config.num_frames=8' \
'model_config.camera_desc_lookup_path=./extracted_meta_with_descriptions.json' \
'model_config.confusion_score_mode=match_binary'See robometer/eval_commands/ for the other baselines. Dataset converters for LIBERO, AgiBotWorld, and custom datasets (DROID / Bridge style) live in robometer/dataset_upload/.
Policy Ranking: Attention Mask and Value Tokenizer
Two knobs change the ranking score without any retraining: the value tokenizer the temporal head was trained with, and the inference-time attention mask. Every row below was trained with value-isolation attention; paired rows differ only in the mask used at inference, on the same checkpoint and the same processed trajectories (max_frames=8, stride=2, num_frames=8, per-trajectory camera descriptions).
Kendall's ฯโ. The last row is the headline configuration reported in Results.
Attention. Value isolation lets queries interact within a prediction group but blocks their representations from later observation and value groups; the causal mask opens that path at inference, so later queries can read earlier value-query states (though not their decoded scalars). On paired trajectories this consistently helps: 0.646 โ 0.704 (8B quantile), 0.614 โ 0.679 (4B quantile), 0.584 โ 0.667 (8B normalized). Earlier value states appear to disambiguate the final queried score by carrying trajectory context, though terminal-score ranking alone cannot establish better frame-level distance estimates or sensitivity to local regression. Note this is not the same intervention as the w/o Isolation ablation in Results, which removes the mask during training and drops the average to 0.611 โ that one permits the cross-group shortcut throughout learning.
Value tokenizer. Frame-level labels are strongly skewed (median 11.57 s, 90th percentile 68.73 s, max 636.33 s), which motivates non-uniform resolution: 119 of 256 quantile centers fall within 10 s and 217 within 50 s, whereas only ~20 of the same 256 centers land within 50 s when spaced at fixed intervals up to the same maximum. Normalized progress changes the target scale altogether โ the same episode fraction can mean different remaining durations, weakening the shared cost-to-go interpretation. On the mean, 8B quantile beats normalized under isolation (0.646 vs. 0.584) and beats both normalized and fixed bins under a causal mask (0.704 vs. 0.667 / 0.675). The suite-level picture is mixed, though: normalized 8B leads on USC Koch, and fixed-bin 4B leads on USC xArm and UTD SO101. Fixed bins may help at long horizons and normalized progress where fraction-completed matters; label density is a plausible account of the mean advantage rather than a demonstrated cause of every per-suite difference.
Reproducing a row. bash rynnvalue_eval/policy_ranking.sh [sdpa|eager|isolated] selects the inference mask: sdpa (the default) is the causal mask applied through Transformers' built-in scaled dot-product attention backend and reproduces the last row of the table above for the released 8B quantile checkpoint; isolated keeps the training-time value-isolation mask (pred_slot_isolated_eager); eager runs the same causal mask through the eager kernel. The script always passes the mask explicitly. If you invoke run_baseline_eval.py directly, do the same โ left unset, the model falls back to whatever its own config.json records, which is not the same across checkpoint generations.
Reward Server
Serve RynnValue as an HTTP reward model (used by policy evaluation and online RL clients):
cd robometer
bash rynnvalue_eval/start_server.sh \
--model-path /path/to/RynnValue-8B \
--port 8001 --gpu 0 --num-frames 8 --batch-size 16The server (robometer/robometer/evals/baseline_eval_server.py) accepts frame sequences + instructions and returns per-frame values / rewards. --checkpoint-path alternatively loads a raw training checkpoint (model.pt with a sibling huggingface/ snapshot). --mode absolute selects the remaining-time head; --debug attaches debugpy on :5678.
Policy RL with pi-rl
pi-rl/ is a fork of openpi that turns value/reward signals into better VLA policies. On top of upstream ฯโ / ฯโ-FAST / ฯโ.โ
SFT (JAX and PyTorch), it adds:
- Offline IQL fine-tuning (
scripts/train_iql.py): each step (1) updates a jaxrl2PixelIQLcritic/value, (2) computes the IQL advantage, (3) updates the ฯโ.โ flow-matching policy with advantage-weighted BC (exp(A_scaling ยท adv)). RL-aware plumbing lives insrc/openpi/training/{rl_data_loader,iql_checkpoints,episode_filter}.py. Registered configs includepi05_robotwin_iql,pi05_franka_single_iql,pi05_franka_dual_iql, and variants. - Online DSRL-style SAC latent steering: a small
PixelSACagent steers the frozen ฯโ.โ policy's latent noise online โ in simulation (examples/dsrl_sim/, LIBEROOffScreenRenderEnv) and on a real Franka over WebSocket (examples/dsrl_franka/). - Filter-BC (filtered behavior cloning): a data-curation baseline that needs no critic and no policy-side conditioning. RynnValue labels every demonstration frame with a
progressscore; frames whose ฮป-discounted progress return over a W-step window does not exceed a per-repo threshold are dropped, and the surviving frames train as plain BC (src/openpi/training/progress_advantage.py, configpi05_franka_mixed_filter_bc). See Filter-BC. - Benchmarks & robots: LIBERO, RoboTwin, ALOHA sim, DROID, and a full real-Franka pipeline (
examples/franka/: serving, fine-tuning, async online LoRA training) with LeRobot-format data converters (scripts/convert_franka_data_to_lerobot.py).
Representative commands:
cd pi-rl
# supervised fine-tuning on RoboTwin
FSDP_DEVICES=2 bash scripts/train_pi05_robotwin.sh adjust_bottle-demo_clean_collect_200-50
# offline IQL on RoboTwin
bash scripts/train_iql_robotwin.sh
# Filter-BC on mixed single-arm + dual-arm Franka repos
bash scripts/train_pi05_franka_mixed_filter_bc.shFilter-BC: Progress-Filtered Fine-Tuning
Filtered behavior cloning uses RynnValue purely as an offline data curator โ no critic, no value function, no policy-side conditioning. Every demonstration frame carries a progress label p_t โ [0, 1]; the per-frame reward is the progress increment r_t = p_{t+1} โ p_t and the filtering score is the ฮป-discounted return over a W-step window,
A_t = ฮฃ_{k=0..Wโ1} ฮป^k ยท r_{t+k}, with r โก 0 past the episode endA frame is kept iff A_t > ฯ_repo (strictly), so under the default "zero" criterion only frames that make some forward progress within the next W steps survive. The kept frames then train as plain BC with the task prompt only. Thresholds and frame whitelists are computed once from the lerobot parquet progress columns and cached under assets/<config>/<asset_id>/, with the criterion, W, quantile and ฮป all encoded into the cache filename so a changed setting never reuses a stale whitelist.
cd pi-rl
# default: W=50, lambda=0.98, criterion=zero over the config's repo_ids
bash scripts/train_pi05_franka_mixed_filter_bc.sh
# your own progress-labeled dumps (norm stats must already exist for this override)
REPO_IDS="/data/close_the_drawer,/data/pick_up_the_box" SKIP_NORM=1 \
bash scripts/train_pi05_franka_mixed_filter_bc.sh
# keep only the top-30% of frames by windowed advantage, over a wider window
ADVANTAGE_HORIZON=128 ADVANTAGE_CRITERION=quantile ADVANTAGE_QUANTILE=0.7 \
bash scripts/train_pi05_franka_mixed_filter_bc.shpi05_franka_mixed_sft is the unfiltered control arm โ same mixed single-arm + dual-arm 16-dim action layout, same prompt, same hyperparameters, but every frame trains. The two runs isolate the effect of progress filtering. pi05_franka_mixed_filter_bc needs repos that carry a progress feature (the RynnValue-labeled dumps); pi05_franka_mixed_sft does not.
On the four real Franka tasks, this filter with its default criterion (ฮป = 0.98, W = 50, keep A_t > 0) lifts multi-task BC success from 35.0% to 42.5% over the unfiltered arm โ see Results for the per-task breakdown.
DSRL Online RL on a Real Franka
examples/dsrl_franka/ runs online SAC latent steering on a real (single- or dual-arm) Franka, optionally shaped by a RynnValue reward server. The launcher template is examples/dsrl_franka/scripts/run_train_franka.sh; the flow is:
1. (Optional) Start the RynnValue reward server for reward shaping (see Reward Server):
cd robometer && bash rynnvalue_eval/start_server.sh --model-path /path/to/RynnValue-8B --port 80002. Start the robot environment server (WebSocket) on the machine controlling the Franka, or skip this and use --fake_env for a quick dry run without hardware.
3. Launch online training. Point it at an IQL-initialized checkpoint (FRANKA_SFT_CKPT_BASE) and the env/reward servers:
cd pi-rl
export FRANKA_SFT_CKPT_BASE=/path/to/iql_checkpoint/10000 # offline-IQL warm start
export XLA_PYTHON_CLIENT_PREALLOCATE=false
python examples/dsrl_franka/launch_train_franka.py \
--arm_mode dual \
--policy_config pi05_franka_dual_iql_optimized_v2 \
--update_type episode --utd_ratio 100 --publish_hz 10.0 \
--client_host localhost --client_port 8101 \
--pi0_action_horizon 16 \
--franka_norm_stats_asset_id pick_up_the_box \
--max_episodes 500 --franka_max_timesteps 600 \
--start_online_updates 200 \
--noise_episodes 2 --noise_std 0.1 \
--score_server "http://localhost:8000" \
--shaping_weight 1.0 --shaping_gamma 0.999 \
--task_description "Move the box from the right side to the left side." \
--checkpoint_interval 200 --checkpoint_dir ./experiments/dual_reward_shaping \
--wandb_project dsrl_franka --seed 42Key knobs:
The same recipe runs in simulation via examples/dsrl_sim/launch_train_sim.py (LIBERO OffScreenRenderEnv).
See pi-rl/README.md for upstream documentation (checkpoints under gs://openpi-assets/checkpoints/: pi05_base, pi05_libero, pi05_droid, โฆ), norm-stat computation, and policy serving.
Checkpoint Conversion
Export a raw training checkpoint to a standalone HuggingFace model directory (embeds the modeling code for trust_remote_code loading):
uv run python tools/convert_rynn_value_lang_to_hf.py \
--model_ckpt_path /path/to/checkpoint_model_XXXXXX \
--output_path /path/to/RynnValue-8B-hf \
--dtype bf16--model_ckpt_path expects a directory containing model.pt with the matching huggingface/ processor/config snapshot next to it.
Model Zoo
RynnValue provides 4B and 8B checkpoints with either fixed-bin or quantile value tokenization. All variants use absolute and relative distributional heads to predict remaining time and signed temporal displacement in seconds, and generate trajectory analysis (Video Description, Match, and Success).
Scores are average Kendall's ฯโ across the six RBM-EVAL-OOD policy-ranking suites, evaluated with a causal attention mask. The headline results in this README correspond to the Quantile variants; see Policy Ranking: Attention Mask and Value Tokenizer for the full comparison.
See the ModelScope collection for mirrored checkpoints. Load any variant with trust_remote_code=True, as in Using RynnValue Programmatically.
Acknowledgements
This repository builds on outstanding open-source work:
- **openpi** (Physical Intelligence) โ ฯโ / ฯโ-FAST / ฯโ.โ
VLA models and training stack (Apache-2.0;
pi-rl/). - **Robometer** โ "Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons": benchmark, RBM baselines, and the RBM-1M dataset (
robometer/). - **jaxrl2** โ IQL / SAC agents (vendored in
pi-rl/third_party/jaxrl2/). - **RynnBrain** - Open Embodied Foundation Models
- **Qwen3-VL** โ the underlying vision-language architecture of the RynnBrain backbone.
- Simulation benchmarks: LIBERO, RoboTwin, DROID, gym-aloha.
License
- The RynnValue components (
rynn_value/,rynn_infer/,tools/) are distributed under the Apache License 2.0 (see `LICENSE`). pi-rl/is distributed under the Apache License 2.0 (seepi-rl/LICENSE) with additional Gemma terms inpi-rl/LICENSE_GEMMA.txt.robometer/follows the upstream Robometer project under the MIT License (seerobometer/LICENSE). It additionally vendors FSDP utilities derived from ByteDance's verl project (Apache-2.0; original copyright headers retained inrobometer/robometer/utils/fsdp/).
Citation
If you find RynnValue useful, please cite our paper:
@article{huang2026rynnvalue,
title = {RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance},
author = {Huang, Dongchi and Zhang, Hongyin and Hou, Bohan and Huang, Siteng and
Su, Zhian and Guo, Hang and Lu, Tong and Tang, Jiahao and Yang, Jianfei and
Wang, Donglin and Peng, Peixi and Chen, Mingxiu and Zhao, Deli and Li, Xin},
journal = {arXiv preprint arXiv:2608.09853},
year = {2026},
doi = {10.48550/arXiv.2608.09853},
url = {https://arxiv.org/abs/2608.09853}
}