World Critic Model Boosts Vision-Language-Action Reinforcement Learning

World Critic Model Boosts Vision-Language-Action Reinforcement Learning

Senyu Fei, Xiaopeng Yu, Siyin Wang, Xianzhong Zhao, Jingjing Gong +1 បន្ថែម

5 នាទី​អាន6 សីហា 2026

Existing critic-based vision-language-action reinforcement learning methods struggle to capture temporal structure in partially observable environments. This work introduces the World Critic Model (WCM), a unified architecture that jointly learns latent state prediction and value estimation. Experiments across 149 tasks in simulation and seven real-world manipulation tasks demonstrate consistent state-of-the-art performance in both in-distribution and out-of-distribution settings.

Experimental Setup

Evaluations were conducted on 7 robotic tasks using the WidowX-250S platform, covering a dynamic grasping task, a long-horizon task, 2 deformable object manipulation tasks, and 3 pick-and-place tasks. Using an off-policy RL algorithm and OpenVLA-OFT as base policies, the models were trained with off-policy reinforcement learning guided by the World Critic Model.

Comparison of in-distribution and out-of-distribution performance across benchmarks

Generalization Performance

Generalization is a key evaluation criterion in VLA-RL. The method was evaluated under the ManiSkill-OOD setting and LIBERO-Plus, yielding two key findings.

Strong generalization gains with WCM. On ManiSkill, WCM improves both in-distribution performance and OOD generalization, outperforming traditional Flow-SDE and PPO. Additionally, the method outperforms the Step-NFT baseline, which is known for its strong OOD performance, benefiting from WCM's ability to capture more state information and provide robust value estimation under distribution shifts.

Superior generalization than SFT. In LIBERO-Plus, starting from one-shot supervised fine-tuning, after approximately 250 RL training steps, the method outperforms full-shot SFT trained on 20k trajectories.

Real-World Performance

The real-world experiments demonstrate the effectiveness and efficiency of WCM under limited-data training on physical robots, operating at the scale of hundreds to a few thousand data samples. With only hundreds to a few thousand trajectories and less than one hour of training, WCM enables rapid iterative refinement and yields accurate critic predictions.

Real-world robot manipulation tasks used for evaluation with the WCM framework

Is Longer Observation History Always Beneficial?

An ablation study varied the observation history length of WCM from 1 to 5 frames. Length 3 performed best on average. A plausible intuitive explanation is that three consecutive frames may implicitly capture second-order dynamics (acceleration), while two frames capture first-order dynamics (velocity). For the tested tasks, first- and second-order information appear sufficient to describe the required dynamic features.

Inference Throughput

Inference throughput was evaluated across three tasks: towel folding, stovetop cleaning, and rotating sushi picking, measured as the number of successful rollouts per hour.

The SFT-only model yields the lowest throughput, due to its low initial success rate and inefficient action execution, characterized by frequent pauses and small-magnitude movements which prolong each rollout. After RL fine-tuning, throughput improves significantly as the policy becomes both more successful and smoother.

Among the two RECAP-trained variants, the model using the WCM-based critic consistently outperforms the Gemma VLM-based critic variant across all tasks, indicating that the choice of critic model has a substantial impact on inference efficiency.

Training Behavior

Training curves were recorded from step 0 until each setting reaches its reported optimal performance across all eight configurations. Most settings exhibit stable improvement over training steps, though convergence speeds and final performance vary across configurations. For ManiSkill-OOD and LIBERO-Plus, since test settings differ from training settings, the reported numerical results show certain discrepancies compared to the training curves; these differences reflect the generalization gap under distribution shift.

Conclusion

This work identified a fundamental limitation of existing critic-based VLA-RL methods: value estimation from single-frame observations or weakly supervised history fails to capture the temporal structure required for state reconstruction under partial observability. The World Critic Model addresses this with a unified architecture that jointly learns latent state prediction and value estimation. Extensive experiments on 149 tasks across four simulation benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. Validation on seven real-world manipulation tasks confirms its stable and effective deployment.

Frequently Asked Questions

What is the World Critic Model (WCM)? WCM is a unified architecture for vision-language-action reinforcement learning that jointly learns latent state prediction and value estimation. It addresses the limitation of single-frame value estimation that fails under partial observability.

How does WCM perform on out-of-distribution tasks? WCM consistently outperforms existing baselines including Flow-SDE, PPO, and Step-NFT on OOD benchmarks. In LIBERO-Plus, it surpasses full-shot SFT trained on 20k trajectories after only 250 RL training steps.

What observation history length works best? An ablation study found that a history length of 3 frames performs best on average. Three consecutive frames implicitly capture second-order dynamics (acceleration), which appears sufficient for the tested manipulation tasks.

How does WCM perform in real-world deployment? WCM enables effective RL training on physical robots with only hundreds to a few thousand trajectories and less than one hour of training. It also improves inference throughput by producing smoother, more successful policies.

🍪 ចំណូលចិត្តខូគី

យើងប្រើខូគីដើម្បីវាស់ប្រសិទ្ធភាព។ គោលនយោបាយ​ឯកជន