TY - GEN
T1 - COMPRESSING HUMAN BODY VIDEO WITH INTERACTIVE SEMANTICS
T2 - 32nd IEEE International Conference on Image Processing, ICIP 2025
AU - Chen, Bolin
AU - Yin, Shanzhi
AU - Zhu, Hanwei
AU - Zhu, Lingyu
AU - Zhang, Zihan
AU - Chen, Jie
AU - Liao, Ru Ling
AU - Wang, Shiqi
AU - Ye, Yan
N1 - Publisher Copyright:
©2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body signal into a series of configurable embeddings, which are controllably edited, compactly compressed, and efficiently transmitted. Moreover, the proposed decoder can evolve the mesh-based motion fields from these decoded semantics to realize the high-quality human body video reconstruction. Experimental results illustrate that the proposed framework can achieve promising compression performance for human body videos at ultra-low bitrate ranges compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes. Furthermore, the proposed framework enables interactive human body video coding without any additional pre-/post-manipulation processes, which is expected to shed light on metaverse-related digital human communication in the future.
AB - In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body signal into a series of configurable embeddings, which are controllably edited, compactly compressed, and efficiently transmitted. Moreover, the proposed decoder can evolve the mesh-based motion fields from these decoded semantics to realize the high-quality human body video reconstruction. Experimental results illustrate that the proposed framework can achieve promising compression performance for human body videos at ultra-low bitrate ranges compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes. Furthermore, the proposed framework enables interactive human body video coding without any additional pre-/post-manipulation processes, which is expected to shed light on metaverse-related digital human communication in the future.
KW - controllable embeddings
KW - deep generative model
KW - human body video
KW - Interactive video coding
UR - https://www.scopus.com/pages/publications/105028586663
U2 - 10.1109/ICIP55913.2025.11084296
DO - 10.1109/ICIP55913.2025.11084296
M3 - 会议稿件
AN - SCOPUS:105028586663
T3 - Proceedings - International Conference on Image Processing, ICIP
SP - 157
EP - 162
BT - 2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PB - IEEE Computer Society
Y2 - 14 September 2025 through 17 September 2025
ER -