Hierarchical Post-Training for Reliable Robotic Manipulation

Hierarchical Post-Training for Reliable Robotic Manipulation

He Kong, Zengjue Chen, Qi Wang, Qianli Xing, Runliang Niu +4 več

7 min branja9. avg. 2026

Methodology

We propose a hierarchical decomposition framework for robotic manipulation control named HiRoC. In HiRoC, the planner and executor are designed as two complementary modules. The planner, trained via supervised fine-tuning (SFT), generates subgoals according to the current situation. These subgoals guide the executor in making decisions, with SFT and reinforcement learning (RL) tuning employed to train the executor.

Training the Planner by SFT

Existing backbones for embodied AI fail to decompose tasks into subtasks and dynamically generate corresponding subgoals. To overcome this issue, we fine-tune our planner via SFT. However, the corresponding data are insufficient for training the planner. Thus, we clean and reorganize existing datasets for planner training.

Training Planner

Based on the prepared trajectories, the planner is optimized via SFT to generate the corresponding subgoal conditioned on the current observation. Specifically, given the reconstructed dataset, the SFT objective is formulated as follows:

The SFT objective trains the planner to predict the appropriate subgoal from the current observation and reconstructed training data.

Experiments

We adopt Qwen2.5-VL-3B as the planner and fine-tune it with LoRA on our prepared data. OpenVLA-OFT is employed as the executor. The executor is first fine-tuned via SFT with LoRA on the reorganized subgoal-conditioned data, followed by reinforcement learning tuning.

The group size of GRPO is set to 8. During evaluation, we perform 50 test episodes for each task and report the average success rate for each task and benchmark suite. All experiments are conducted on 8 NVIDIA H200 GPUs. We compare HiRoC with ten representative baselines. More details are provided in the Appendix.

Transfer Ability: Zero-Shot Performance

To evaluate the generalization capability of HiRoC, we validate it on LIBERO-Plus, where OpenVLA, OpenVAL*-One, and WorldVLA are compared. The success rate of each perturbation and the average success rate are reported.

HiRoC achieves the best generalization across all perturbation types. Its substantial improvement over SFT-based models such as OpenVLA indicates that online interaction explores more diverse trajectories than a static full-shot dataset.

Although WorldVLA improves generalization by modeling environment dynamics, it remains sensitive to environmental changes and requires additional data to adapt to new tasks. In contrast, HiRoC decomposes the goal into subgoals, allowing the executor to focus on each subgoal and thereby reducing its burden.

Zero-shot transfer evaluation across LIBERO-Plus perturbation types

Real-World Experiments

To validate the practical applicability of HiRoC, we conduct a sim-to-real experiment on a real robotic platform. During training, the executor is optimized entirely in a high-fidelity simulator, where the workspace, object dimensions, and robot configuration are consistent with those of the real system.

The trained VLA policy is then directly deployed to the real robot without additional fine-tuning. Guided by the planner-generated subgoals, the executor successfully approaches, grasps, transports, and places the target object.

The successful deployment demonstrates the effectiveness of HiRoC for real-world robotic manipulation and its promising sim-to-real transfer capability.

Sim-to-real deployment of hierarchical robotic manipulation

Conclusion

We propose HiRoC, a hierarchical post-training framework with a high-level planner and a low-level executor. The planner, trained via SFT, decomposes complex tasks into simpler subtasks and generates subgoals to guide execution.

To address the distribution misalignment between the two modules, we reorganize the subgoal-conditioned dataset and further optimize the executor through RL tuning. Experimental results and extensive analyses demonstrate the effectiveness of HiRoC and each of its components.

Although HiRoC demonstrates the value of hierarchical decomposition for robotic manipulation, future work will explore end-to-end training and more robust decomposition mechanisms for handling uncertainty in real-world physical systems. We hope HiRoC underscores the importance of combining hierarchical task decomposition with reinforcement learning-based policy optimization for complex robotic manipulation.

Experimental Details

Simulation Experiments

We conduct all experiments on the four LIBERO benchmark suites, including Spatial, Object, Goal, and the long-horizon LIBERO-10 suite.

Following prior work, we first fine-tune the pretrained OpenVLA-OFT model on a small subset of demonstrations from each suite to obtain a suite-specific base executor. The high-level planner is initialized from Qwen2.5-VL-3B and fine-tuned via SFT on our curated planner dataset to generate intermediate language subgoals.

Real-World Experiments

We conduct the real-world experiments in JoySim using JoyRA-0.1 as the base VLA model. The task requires the robot to grasp a correction fluid bottle from the tabletop and place it into a box.

The high-level planner is initialized from Qwen2.5-VL-3B and fine-tuned via SFT, while the low-level executor is initialized from JoyRA-0.1 and optimized using Flow-SDE through reinforcement learning in simulation.

After training, the learned policy is directly deployed to the real robotic platform for zero-shot evaluation without any additional real-world fine-tuning.

Prompt Details

The prompts for the high-level planner and low-level executor are provided in the paper.

Case Study

Several representative manipulation cases illustrate the effectiveness of HiRoC in long-horizon robotic tasks. Unlike flat VLA policies that continuously condition on the global instruction, HiRoC explicitly decomposes complex tasks into sequential subgoals through the high-level planner, allowing the executor to focus on the current manipulation objective.

In one case, the robot is required to pick up the cream cheese and place it into the basket. The planner first generates the subgoal “approach cream cheese,” guiding the executor to locate and reach the target object. After the object is approached, the planner updates the objective to “lift cream cheese,” enabling the executor to perform accurate grasping actions.

Subsequently, the subgoal is changed to “move to basket” and “place in basket,” respectively. By providing stage-specific semantic guidance, HiRoC avoids the ambiguity of directly executing the complete task instruction and successfully completes the long-horizon manipulation sequence.

A second case demonstrates the ability of HiRoC to handle tasks requiring sequential interactions with different spatial targets. For the instruction of placing the bowl into the bottom drawer, the planner gradually decomposes the task into approaching, lifting, and placing stages.

Instead of requiring the executor to infer the entire manipulation procedure from the global instruction, each subgoal provides a clear intermediate objective. Consequently, the executor can adapt its behavior according to the current task stage and maintain consistent progress throughout the execution process, showing the advantage of hierarchical task decomposition.

A third case illustrates a more complex multi-object manipulation scenario, where the robot needs to place two moka pots onto the stove. The planner first guides the executor to approach and lift the first pot, followed by placing it on the stove, and then generates new subgoals for manipulating the second pot.

This case highlights that HiRoC can dynamically organize multiple manipulation steps and handle tasks with repeated object-level operations. By combining planner-generated subgoals with subgoal-conditioned executor optimization, HiRoC provides reliable guidance for long-horizon decision making and achieves robust task completion.

Frequently Asked Questions

What is HiRoC? HiRoC is a hierarchical post-training framework with a high-level planner and a low-level executor for robotic manipulation.

How is the planner trained? The planner is initialized from Qwen2.5-VL-3B and fine-tuned with supervised fine-tuning to generate intermediate language subgoals.

How is the executor optimized? The executor is first trained with supervised fine-tuning on reorganized subgoal-conditioned data and then optimized through reinforcement learning tuning.

Does HiRoC transfer to a real robot? Yes. A policy trained in simulation was directly deployed to a real robotic platform for zero-shot evaluation without additional real-world fine-tuning.

🍪 Nastavitve piškotkov

Uporabljamo piškotke za merjenje zmogljivosti. Politika zasebnosti