GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao, Yi-chung Chen +7 lagi

3 min baca25 Ogo 2026

World Models for Autonomous Driving

Beyond video generation, occupancy-based world models forecast scene evolution in a voxelized three-dimensional space. OccWorld tokenizes 3D occupancy and autoregressively predicts future occupancy together with ego motion, while Drive-OccWorld extends occupancy forecasting with action conditioning and occupancy-based planning.

These approaches are supervised with voxelized ground-truth occupancy targets, requiring occupancy annotations to construct their prediction space. In contrast, GeoWAM does not require ground-truth occupancy annotations and learns future geometry from dense metric point-map targets derived from off-the-shelf geometry foundation models, requiring only RGB images for training.

GeoWAM world action model pipeline

Visual Geometry Models

Driving scenes introduce additional requirements, including long-range metric accuracy, temporal motion, surround-view observations, and substantial variation across camera configurations. DVGT addresses these requirements with an ego-centric formulation that predicts metric point maps and ego poses from multi-frame, multi-view images using factorized spatial-temporal attention, while DVGT-2 scales this geometry representation toward autonomous-driving action modeling.

However, existing visual geometry models primarily reconstruct the scene contained in observed images. GeoWAM extends visual geometry learning from reconstruction to forecasting, using future point-map prediction to learn scene dynamics and connect geometric representations with trajectory planning.

GeoWAM visual geometry representation

Planning on NAVSIM V2

Table 2 compares GeoWAM with perception-based planners and recent world-action models on the navtest split. GeoWAM achieves an EPDMS of 90.2, improving upon the DVGT-2 initialization by 0.6 points and establishing the best overall score in the table.

It also matches the best DDC and TLC scores, while remaining competitive across the other safety and progress components. Together, these results demonstrate the effectiveness of the visual geometry formulation with a deterministic trajectory decoder.

Conclusion

We presented GeoWAM, a visual geometry world action model for autonomous driving. Instead of modeling scene evolution through future-image generation, GeoWAM learns to forecast future geometric features from historical multiview observations under both feature-level and dense point-map supervision.

Following geometry pretraining, it adopts an inverse-dynamics-like formulation that infers future ego tokens from the predicted geometric dynamics and maps them to an ego trajectory with a geometry-conditioned action head. This two-stage design transfers predictive geometric knowledge directly to planning without requiring future image synthesis.

By placing geometric evolution between observation and action, GeoWAM provides a spatially grounded formulation for world action modeling in autonomous driving.

Frequently Asked Questions

What is GeoWAM? GeoWAM is a visual geometry world action model for autonomous driving.

What supervision does GeoWAM use? It uses feature-level and dense point-map supervision derived from geometry foundation models and requires only RGB images for training.

How does GeoWAM connect geometry prediction with planning? It infers future ego tokens from predicted geometric dynamics and maps them to an ego trajectory with a geometry-conditioned action head.

How does GeoWAM perform on the NAVSIM V2 navtest split? It achieves an EPDMS of 90.2, improving upon the DVGT-2 initialization by 0.6 points.

🍪 Keutamaan kuki

Kami menggunakan kuki untuk mengukur prestasi. Dasar Privasi