Skip to main navigation Skip to search Skip to main content

Standardizing Generative Face Video Compression using Supplemental Enhancement Information

  • Bolin Chen
  • , Yan Ye*
  • , Jie Chen
  • , Ru Ling Liao
  • , Shanzhi Yin
  • , Shiqi Wang
  • , Kaifa Yang
  • , Yue Li
  • , Yiling Xu
  • , Ye Kui Wang
  • , Shiv Gehlot
  • , Guan Ming Su
  • , Peng Yin
  • , Sean McCarthy
  • , Gary J. Sullivan
  • *Corresponding author for this work
  • Alibaba Group Holding Ltd.
  • City University of Hong Kong
  • Shanghai Jiao Tong University
  • ByteDance Ltd.
  • Dolby Laboratories

Research output: Contribution to journalArticlepeer-review

Abstract

This paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D key-points, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274 | ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Generative AI
  • JVET
  • VSEI
  • VVC
  • face video coding
  • supplemental enhancement information

Fingerprint

Dive into the research topics of 'Standardizing Generative Face Video Compression using Supplemental Enhancement Information'. Together they form a unique fingerprint.

Cite this