{OC}-{SOP}: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
Autonomous driving perception faces significant challenges due to occlusions and incomplete scene data in the environment. To overcome these issues, the task of semantic occupancy prediction ({SOP}) is proposed, which aims to jointly infer both the geometry and semantic labels of a scene from images. However, conventional camera-based methods typically treat all categories equally and primarily rely on local features, leading to suboptimal predictions, especially for dynamic foreground objects. To address this, we propose Object-Centric {SOP} ({OC}-{SOP}), a framework that integrates high-level object-centric cues extracted via a detection branch into the semantic occupancy prediction pipeline. This object-centric integration significantly enhances the prediction accuracy for foreground objects and achieves state-of-the-art performance among all categories on {SemanticKITTI}.
- Published in:
2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) - Type:
Inproceedings - Authors:
- Year:
2025 - Source:
https://ieeexplore.ieee.org/document/11343108
Citation information
: {OC}-{SOP}: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness, 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2025, 6911--6918, October, https://ieeexplore.ieee.org/document/11343108, Cao.Behnke.2025a,
@Inproceedings{Cao.Behnke.2025a,
author={Cao, Helin; Behnke, Sven},
title={{OC}-{SOP}: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness},
booktitle={2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)},
pages={6911--6918},
month={October},
url={https://ieeexplore.ieee.org/document/11343108},
year={2025},
abstract={Autonomous driving perception faces significant challenges due to occlusions and incomplete scene data in the environment. To overcome these issues, the task of semantic occupancy prediction ({SOP}) is proposed, which aims to jointly infer both the geometry and semantic labels of a scene from images. However, conventional camera-based methods typically treat all categories equally and primarily...}}