TY - JOUR
T1 - Space–Time Gaussian Surfels for High-Fidelity Dynamic Objects Segmentation and Representation
AU - Zheng, Xiaoyun
AU - Li, Xufeng
AU - Liao, Liwei
AU - Gao, Feng
AU - Wang, Shiqi
AU - Wang, Ronggang
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026/5/1
Y1 - 2026/5/1
N2 - We introduce ST-ObjGS, a method using Space-time Gaussian surfels for accurate object segmentation within 4D representations. Our approach addresses the limitations of current Gaussian-based methods, which primarily focus on static 3D scene understanding and struggle with geometrically accurate object segmentation in complex dynamic scenes. To ensure robust object-level segmentation, we first integrate Grounded SAM 2, which enables text prompt-based object selection and tracking. We then learn a set of Gaussian surfels for object geometry representation and employ a marginal 1D Gaussian for dynamic modeling at each timestamp. To improve geometric quality when modeling surfaces, we use depth and surface normal for geometric regularization. Furthermore, to address continuity and flickering issues in complex scenes, we implement dynamic-aware regularization to maintain temporal consistency. This approach allows us to capture object motion and morphing over time while maintaining spatial coherence. To the best of our knowledge, ST-ObjGS is the first self-supervised approach using Space-time Gaussian surfels for consistent segmentation of dynamic 3D objects in real-world scenes. Extensive experiments on standard benchmarks including PKU-DyMVHumans, Plenoptic Video, Google Immersive, and CMU Panoptic datasets demonstrate that ST-ObjGS produces more precise object masks than its Gaussian-based counterparts and significantly outperforms supervised single-view baselines.
AB - We introduce ST-ObjGS, a method using Space-time Gaussian surfels for accurate object segmentation within 4D representations. Our approach addresses the limitations of current Gaussian-based methods, which primarily focus on static 3D scene understanding and struggle with geometrically accurate object segmentation in complex dynamic scenes. To ensure robust object-level segmentation, we first integrate Grounded SAM 2, which enables text prompt-based object selection and tracking. We then learn a set of Gaussian surfels for object geometry representation and employ a marginal 1D Gaussian for dynamic modeling at each timestamp. To improve geometric quality when modeling surfaces, we use depth and surface normal for geometric regularization. Furthermore, to address continuity and flickering issues in complex scenes, we implement dynamic-aware regularization to maintain temporal consistency. This approach allows us to capture object motion and morphing over time while maintaining spatial coherence. To the best of our knowledge, ST-ObjGS is the first self-supervised approach using Space-time Gaussian surfels for consistent segmentation of dynamic 3D objects in real-world scenes. Extensive experiments on standard benchmarks including PKU-DyMVHumans, Plenoptic Video, Google Immersive, and CMU Panoptic datasets demonstrate that ST-ObjGS produces more precise object masks than its Gaussian-based counterparts and significantly outperforms supervised single-view baselines.
KW - 3D Gaussian
KW - Gaussian surfels representation
KW - dynamic object segmentation
UR - https://www.scopus.com/pages/publications/105024657485
U2 - 10.1109/TCSVT.2025.3642694
DO - 10.1109/TCSVT.2025.3642694
M3 - 文章
AN - SCOPUS:105024657485
SN - 1051-8215
VL - 36
SP - 5747
EP - 5758
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
IS - 5
ER -