One Policy Controls Three Robots Using Shared Skills Across Different Bodies

One Policy Controls Three Robots Using Shared Skills Across Different Bodies

Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng +10 aktar

7 min ta' qari7 ta’ Aww, 2026

One DyPES-VLA checkpoint reaches 89.02% on RoboTwin, 59.25% on a humanoid benchmark, and 98.0% on LIBERO, despite the robots using different bodies and action spaces. The system combines video-based future prediction with embodiment-specific control heads, giving one policy a shared understanding of manipulation while preserving the low-level commands each robot needs.

What Did the Researchers Build?

DyPES-VLA is a vision-language-action system designed to control different robot embodiments from a single trained checkpoint. An embodiment is the physical form and control interface of a robot, such as a dual-arm machine, a humanoid, or a single-arm manipulator. These platforms differ in joint counts, movement limits, gripper layouts, and the format of the actions sent to their motors.

The system separates the parts of manipulation that can be shared from the parts that must remain robot-specific. A shared visual representation learns general information about how objects, hands, arms, and environments change over time. Separate action components then translate that shared information into commands appropriate for each robot.

The training process has two stages. First, the system learns from large collections of videos without action labels. Those videos include egocentric human manipulation footage from EgoDex and simulated demonstrations generated for the robot embodiments used during co-training. The first stage trains the vision-language model, query tokens that extract task-relevant information, and a SANA head for predicting future visual states.

The second stage uses the learned representation for policy training across multiple embodiments. An embodiment-specific mixture-of-experts action head handles each robot’s native control space, while embodiment metadata tells the system which robot and data source produced a particular input.

Cross-embodiment robot policy architecture and training pipeline

What Were the Key Results?

A single DyPES-VLA checkpoint matched or exceeded the strongest listed methods across three simulation benchmarks. The tests covered a 14-degree-of-freedom dual-arm robot, a 29-degree-of-freedom humanoid, and a 7-degree-of-freedom single-arm robot.

BenchmarkRobot embodimentDyPES-VLA resultComparison
RoboTwin 2.0, clean14-DoF dual-arm robot88.78%2.68 points above Qwen-VLA
RoboTwin 2.0, randomized14-DoF dual-arm robot89.26%2.06 points above Qwen-VLA
RoboTwin 2.0, average14-DoF dual-arm robot89.02%2.37 points above Qwen-VLA
RoboCasa-GR129-DoF humanoid59.25%Above ABot-M0 at 58.3% and Qwen-VLA at 56.7%
LIBERO7-DoF single-arm robot98.0%Above Fast-WAM at 97.6% and OpenVLA-OFT at 97.1%

The largest ablation effect came from removing future prediction. Performance fell by 2.4 points on RoboTwin 2.0 and 2.5 points on RoboCasa-GR1. Removing only the first-stage pretraining also reduced results across all three benchmarks, showing that action-free video learning contributed more than simple initialization.

On LIBERO, DyPES-VLA trailed the best fine-tuned X-VLA result by only 0.1 points. That result is important because DyPES-VLA uses one cross-embodiment checkpoint rather than a separate specialist for each benchmark.

How Does DyPES-VLA Work?

DyPES-VLA addresses two different learning problems: understanding manipulation dynamics and producing valid commands for a particular robot. The shared part focuses on the first problem. The embodiment-specific action head focuses on the second.

During Stage 1, the model receives video sequences and learns a future-prediction objective. It does not need action labels for this step. Instead of only recognizing the current frame, the model is trained to represent how a scene is likely to evolve after an interaction. For manipulation, that information can include whether an object is being grasped, how a tool is moving, or how a hand-object relationship is changing.

The training data combines human egocentric videos with simulated videos from the robot embodiments used in co-training. Human footage provides broad examples of manipulation behavior, while simulation supplies data tied to the target robots and their environments. The vision-language model, query tokens, and SANA head are pretrained together so the resulting representation captures task-relevant visual and temporal information.

During Stage 2, the representation supports action learning across the different robot datasets. A mixture-of-experts, or MoE, action head assigns control prediction to an expert associated with the relevant embodiment. This prevents a 7-DoF arm’s action format from being treated as interchangeable with a 29-DoF humanoid’s format.

Embodiment metadata provides another signal. It identifies the robot or data source associated with an observation, helping the model distinguish between similar visual situations that require different motor outputs. The shared representation therefore handles common manipulation knowledge, while the selected expert converts that knowledge into the robot’s native action space.

Robot embodiments evaluated by the shared manipulation policy

Why Does This Matter for Robotics?

Most robot learning systems face a difficult trade-off. A specialist policy can exploit the details of one robot, but it does not transfer easily to another platform. A generalist policy can share data and capabilities across a fleet, but heterogeneous bodies and action spaces can cause interference: the same visual situation can require completely different joint commands.

DyPES-VLA offers a practical architecture for handling that trade-off. A manufacturer or robotics operator could retain one shared model for perception and manipulation knowledge while attaching control experts for each platform. That approach could reduce duplicated training and make data from different robots more useful.

The results are especially relevant to fleets that mix arms, humanoids, and mobile manipulation systems. Buyers evaluating browse humanoid robots on BotMarket can view the humanoid as one embodiment in a broader policy ecosystem rather than an isolated platform. Similarly, operators comparing used cobots for sale may benefit from policies that preserve shared skills while adapting to different joint layouts and controller interfaces.

The work also highlights the value of video data. Action-free human footage is easier to collect at scale than carefully synchronized robot demonstrations, making future prediction a potentially efficient route to broader manipulation knowledge.

What Are the Limitations and Open Questions?

The reported results come from simulation benchmarks, so they do not establish how well DyPES-VLA handles real sensors, contact uncertainty, calibration errors, hardware wear, or unexpected objects. Simulation also provides cleaner embodiment metadata and more consistent action interfaces than many production environments.

The supplied results do not report inference latency, memory requirements, training cost, or the amount of data needed for each embodiment. Those factors will determine whether a shared checkpoint is practical on constrained robot computers.

Future work also needs to test whether the same architecture transfers to a new robot with limited data, rather than only supporting embodiments included during co-training. Another open question is how many embodiment-specific experts can be added before model management and interference become difficult.

Frequently Asked Questions

What is cross-embodiment manipulation?

Cross-embodiment manipulation means learning manipulation skills that work across robots with different bodies, joint counts, and action interfaces.

What makes DyPES-VLA different from a standard robot policy?

It combines shared future-prediction learning with embodiment-specific action experts instead of forcing every robot to use the same control output format.

Why use human videos without action labels?

Future prediction lets the model learn how manipulation scenes change without requiring motor commands for every video frame, expanding the available training data.

Was DyPES-VLA tested on physical robots?

The reported comparisons cover RoboTwin 2.0, RoboCasa-GR1, and LIBERO simulation benchmarks; the supplied results do not establish physical-robot performance.

What Is the Bottom Line?

DyPES-VLA shows that one policy can perform strongly across robots with very different bodies when shared dynamics learning is separated from embodiment-specific control. Its strongest evidence is the combination of high benchmark scores and the clear performance loss after removing future supervision.

🍪 Preferenzi tal-cookie

Nużaw cookies biex inkejlu l-prestazzjoni. Politika tal-privatezza