A Smarter Way to Build Compact Semantic Maps for Mobile Robots

A Smarter Way to Build Compact Semantic Maps for Mobile Robots

QiYing Deng, ZhongLai Wang, Yuan Gao, Wei Dong

8 мин. четене12.08.2026 г.

M2-SMap gives robots a more compact way to represent 3D scenes without gluing neighboring objects together. The system combines instance-level semantics with three geometric building blocks—bounded planes, superquadrics, and Gaussian mixture model primitives—to reduce map size, preserve object boundaries, and run at 29.37–97.67 frames per second across three RGB-D sequences.

What Did the Researchers Build?

M2-SMap is a memory-efficient semantic mapping system for robots that use RGB-D cameras. An RGB-D camera captures both color and depth, allowing a robot to build a three-dimensional representation of its surroundings. The system is designed to keep that representation compact while retaining meaningful information about surfaces, objects, and the boundaries between them.

The main problem is that compact geometric maps can make poor decisions about how to divide a scene. A geometry-only system might allocate too many small shapes, wasting memory, or combine nearby objects because their surfaces happen to touch. M2-SMap addresses both issues by using semantic information—knowledge about individual object instances—alongside geometric structure.

The resulting map uses three types of representation. Large flat regions become bounded planes, such as finite wall, floor, or tabletop sections. Recognized object regions are modeled with superquadrics, a flexible family of smooth shapes that can approximate forms such as boxes, cylinders, and rounded objects. Remaining details are stored as Gaussian mixture model, or GMM, primitives: compact 3D probability-like blobs that preserve geometry when a simpler shape does not fit.

This design gives the mapper a hierarchy. Broad surfaces receive broad geometric descriptions, object regions receive object-level models, and irregular leftovers retain smaller primitives. The hierarchy reduces unnecessary map elements without forcing every scene region into the same shape.

Hierarchical multi-model representation for compact semantic mapping

What Were the Key Results?

M2-SMap was evaluated on three RGB-D sequences and compared with existing compact multi-model mapping methods. The reported results show improvements in both map compactness and separation between neighboring objects. The strongest result is an 18.7% reduction in average primitive count compared with the best baseline.

A primitive is one geometric element in the map. Fewer primitives generally mean less memory use, simpler map processing, and fewer elements for downstream systems to inspect. The result does not mean that the map discards all fine detail; instead, it indicates that the system represents scene structure more efficiently.

The evaluation also measured inter-object adhesion, a failure in which separate objects become connected in the reconstructed map. M2-SMap reduced the reported adhesion measure from 2.808 to 0. Across the tested sequences, the system maintained processing rates from 29.37 to 97.67 Hz, which the researchers classify as real-time operation.

Evaluation measureM2-SMap resultWhy it matters
Average primitive count18.7% lower than the best baselineMore compact scene representation
Measured inter-object adhesionReduced from 2.808 to 0Better separation of neighboring objects
Processing rate29.37–97.67 HzReal-time mapping across the tested sequences
Evaluation scopeThree RGB-D sequencesDemonstrates performance, but at limited scale

These results point to a practical trade-off improvement: M2-SMap uses fewer map elements while preserving semantic boundaries more effectively than the compared methods.

Comparison of compact geometric representations and object separation

How Does M2-SMap Work?

M2-SMap processes each registered RGB-D frame through a sequence of geometric and semantic decisions.

  1. The camera frame becomes a point cloud.

The depth image supplies 3D positions, while registration aligns the incoming frame with the existing scene representation. The point cloud is then sampled to control the amount of data entering later stages.

  1. The point cloud is divided into supervoxels.

A supervoxel is a small group of nearby 3D points with similar local structure. These groups preserve surface continuity better than treating every point independently and reduce the effect of noisy depth measurements.

  1. Features are examined at two connected scales.

M2-SMap fuses geometric groups to expose larger planar structures. This scale helps determine whether a region should become a flat-plane model. At the same time, the system preserves a mapping back to the original Gaussian scale, so semantic objects and smaller geometric details are not lost during fusion.

  1. Each region receives an adaptive model.

Regions judged to support a flat surface are reconstructed as bounded planes rather than infinite mathematical planes. Semantic object regions are fitted with superquadrics. If a valid superquadric cannot represent an object region, the system falls back to its original GMM components instead of forcing an inaccurate shape.

  1. Residual geometry remains available.

Components that do not belong to a reliable plane or object-level model remain represented by GMM primitives. This fallback is important because real scenes contain irregular surfaces, incomplete observations, and shapes that do not match a simple template.

The method is therefore not a single compression step. It is a model-selection pipeline that chooses the representation appropriate to each part of the scene. Large, simple structures are compressed aggressively; semantic objects receive structured models; and difficult regions retain a more general representation.

The key engineering idea is the link between the two scales. Geometric fusion provides enough context to recognize broad planes, while the preserved fine-scale representation protects object identity and residual detail. That combination directly targets the two weaknesses identified in compact multi-model mapping: inefficient primitive allocation and adhesion between separate objects.

Why Does This Matter for Robotics?

A robot’s map is not just a visual record. Navigation, collision checking, object interaction, and scene understanding all depend on how efficiently and accurately the environment is represented. A map with unnecessary primitives consumes memory and increases the amount of data that later systems must process. A map that merges two objects can create incorrect collision boundaries or make manipulation targets harder to distinguish.

M2-SMap’s 18.7% reduction in average primitive count is relevant to robots operating under limited compute or memory budgets. The reduction can make persistent mapping more practical on embedded hardware, although the reported results do not include a hardware-specific memory profile. Its removal of measured inter-object adhesion is particularly important for cluttered workspaces, where nearby items often have touching or overlapping surfaces in the camera view.

The approach could support warehouse robots that maintain maps around shelving, pallets, containers, and work zones. It also fits the needs of used industrial robots being integrated into changing cells, where compact scene models can help inspection or planning software update its understanding of the workspace.

The result is not a complete navigation or manipulation system. It is a mapping component that gives those systems a cleaner, more compact geometric foundation.

Semantic map components for robot perception and scene understanding

What Are the Limitations and Open Questions?

The evaluation covers three RGB-D sequences, which is enough to demonstrate the reported improvements but not enough to establish performance across every robot, sensor, or environment. The results also provide processing rates and primitive-count reductions without a separate memory-consumption measurement, so the practical memory savings are inferred from the smaller representation rather than quantified directly.

The method depends on registered RGB-D data and usable instance-level semantic information. Errors in alignment, depth sensing, or object assignment could affect the choice between planes, superquadrics, and GMM primitives. The researchers identify larger-scale and dynamic environments as future targets, and tighter integration with robot navigation and manipulation remains open. Those tests will determine whether cleaner maps translate into better task performance.

What Questions Do Buyers and Engineers Ask?

What problem does M2-SMap solve?

It reduces the number of geometric elements needed for a semantic 3D map while preventing separate objects from being merged together.

What is a map primitive?

A map primitive is one compact geometric element used to describe part of a scene, such as a bounded plane, superquadric, or GMM component.

Why combine planes, superquadrics, and GMM primitives?

Each model fits a different type of geometry: broad flat surfaces, recognizable object shapes, and irregular residual details.

Is M2-SMap ready for dynamic warehouse deployment?

The reported tests demonstrate real-time performance on three RGB-D sequences, but larger and dynamic environments remain future evaluation targets.

What Is the Bottom Line?

M2-SMap shows that semantic information can make compact robotic maps both smaller and better at preserving object boundaries. Its combination of hierarchical geometric refinement and adaptive multi-model fitting achieved an 18.7% primitive reduction, eliminated the measured adhesion cases, and maintained real-time processing in the reported tests.

🍪 Предпочитания за бисквитки

Използваме бисквитки за измерване на представянето. Политика за поверителност