TY - GEN
T1 - Light-Weighted Temporal Evolution Inference for Generative Face Video Compression
AU - Zhang, Zihan
AU - Chen, Bolin
AU - Yin, Shanzhi
AU - Wang, Shiqi
AU - Ye, Yan
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Recently, Generative Face Video Compression (GFVC) has advanced the concept of Model-based Coding (MBC) with promising rate-distortion performance relying on the strong inference capabilities of deep generative models. In particular, GFVC can capture temporal evolution of face video using compact representations (i.e., 2D/3D key-points, facial semantics, compact feature), thus achieving the quality and bandwidth trade-offs for ultra-low bit-rate communication. However, there remains an unaddressed challenge, i.e., the existing GFVC models are not light-weighted and low-latency enough for practical applications. To address these obstacles, this paper proposes a practical lightweight scheme based on the Compact Feature Temporal Evolution (CFTE) model, which aims to provide insights into practical deployments and efficient inference. Specifically, the lightweight network architecture is built with depth-wise convolutions and Inverted Residual Blocks to lower the computational complexity. Moreover, a feature-level knowledge distillation is further introduced to improve the performance of lightweight student CFTE model. Experimental results demonstrate that our proposed lightweight GFVC model can achieve an obvious complexity reduction, whilst maintaining competitive rate-distortion performance.
AB - Recently, Generative Face Video Compression (GFVC) has advanced the concept of Model-based Coding (MBC) with promising rate-distortion performance relying on the strong inference capabilities of deep generative models. In particular, GFVC can capture temporal evolution of face video using compact representations (i.e., 2D/3D key-points, facial semantics, compact feature), thus achieving the quality and bandwidth trade-offs for ultra-low bit-rate communication. However, there remains an unaddressed challenge, i.e., the existing GFVC models are not light-weighted and low-latency enough for practical applications. To address these obstacles, this paper proposes a practical lightweight scheme based on the Compact Feature Temporal Evolution (CFTE) model, which aims to provide insights into practical deployments and efficient inference. Specifically, the lightweight network architecture is built with depth-wise convolutions and Inverted Residual Blocks to lower the computational complexity. Moreover, a feature-level knowledge distillation is further introduced to improve the performance of lightweight student CFTE model. Experimental results demonstrate that our proposed lightweight GFVC model can achieve an obvious complexity reduction, whilst maintaining competitive rate-distortion performance.
KW - Generative coding
KW - efficient inference
KW - lightweighted model
UR - https://www.scopus.com/pages/publications/85211317815
U2 - 10.1109/MMSP61759.2024.10743340
DO - 10.1109/MMSP61759.2024.10743340
M3 - 会议稿件
AN - SCOPUS:85211317815
T3 - 2024 IEEE 26th International Workshop on Multimedia Signal Processing, MMSP 2024
BT - 2024 IEEE 26th International Workshop on Multimedia Signal Processing, MMSP 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 26th IEEE International Workshop on Multimedia Signal Processing, MMSP 2024
Y2 - 2 October 2024 through 4 October 2024
ER -