Skip to main navigation Skip to search Skip to main content

Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens

  • Bolin Chen*
  • , Shanzhi Yin*
  • , Zihan Zhang*
  • , Jie Chen
  • , Ru Ling Liao
  • , Lingyu Zhu
  • , Shiqi Wang
  • , Yan Ye
  • *Corresponding author for this work
  • City University of Hong Kong
  • Alibaba Group Holding Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recently, deep generative models have greatly advanced the progress of face video coding towards promising rate-distortion performance and diverse application functionalities. Beyond traditional hybrid video coding paradigms, Generative Face Video Compression (GFVC) relying on the strong capabilities of deep generative models and the philosophy of early Model-Based Coding (MBC) can facilitate the compact representation and realistic reconstruction of visual face signal, thus achieving ultra-low bitrate face video communication. However, these GFVC algorithms are sometimes faced with unstable reconstruction quality and limited bitrate ranges. To address these problems, this paper proposes a novel Progressive Face Video Compression framework, namely PFVC, that utilizes adaptive visual tokens to realize exceptional trade-offs between reconstruction robustness and bandwidth intelligence. In particular, the encoder of the proposed PFVC projects the high-dimensional face signal into adaptive visual tokens in a progressive manner, whilst the decoder can further reconstruct these adaptive visual tokens for motion estimation and signal synthesis with different granularity levels. Experimental results demonstrate that the proposed PFVC framework can achieve better coding flexibility and superior rate-distortion performance in comparison with the latest Versatile Video Coding (VVC) codec and the state-of-the-art GFVC algorithms. The project page can be found at https://github.com/Berlin0610/PFVC.

Original languageEnglish
Title of host publicationProceedings - DCC 2025
Subtitle of host publication2025 Data Compression Conference
EditorsAli Bilgin, James E. Fowler, Joan Serra-Sagrista, Yan Ye, James A. Storer
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages163-172
Number of pages10
ISBN (Electronic)9798331534714
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 Data Compression Conference, DCC 2025 - Snowbird, United States
Duration: 18 Mar 202521 Mar 2025

Publication series

NameData Compression Conference Proceedings
ISSN (Print)1068-0314

Conference

Conference2025 Data Compression Conference, DCC 2025
Country/TerritoryUnited States
CitySnowbird
Period18/03/2521/03/25

Fingerprint

Dive into the research topics of 'Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens'. Together they form a unique fingerprint.

Cite this