Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization
Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing methods typically rely on obtaining rich semantic cues from {RGB} images, which may expose privacy-sensitive visual information. Depth-only 3D geometry provides a privacy-preserving alternative, but the absence of appearance-based semantic cues makes open-vocabulary predictions highly uncertain and less reliable. Under this setting, we propose to convert uncertainty into a guidance signal to identify unreliable semantic responses and use semantic priors from foundation models to regularize their refinement. We present {UTTO}, an uncertainty-guided test-time optimization framework for depth-only open-vocabulary 3D semantic segmentation. Without additional training, experiments on {ScanNet}20, {ScanNet}40, and {ScanNet}200 demonstrate that {UTTO} consistently improves depth-only open-vocabulary 3D segmentation and outperforms representative baselines under privacy-preserving conditions.
- Published in:
arXiv - Type:
Article - Authors:
- Year:
2026 - Source:
http://arxiv.org/abs/2607.00978
Citation information
: Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization, arXiv, 2026, {arXiv}:2607.00978, July, {arXiv}, http://arxiv.org/abs/2607.00978, Huang.etal.2026b,
@Article{Huang.etal.2026b,
author={Huang, Xuying; Pan, Sicong; Bennewitz, Maren},
title={Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization},
journal={arXiv},
number={{arXiv}:2607.00978},
month={July},
publisher={{arXiv}},
url={http://arxiv.org/abs/2607.00978},
year={2026},
abstract={Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing methods typically rely on obtaining rich semantic cues from {RGB} images, which may expose privacy-sensitive visual information. Depth-only 3D geometry provides a...}}