Skip to main navigation Skip to search Skip to main content

COMPRESSING HUMAN BODY VIDEO WITH INTERACTIVE SEMANTICS: A GENERATIVE APPROACH

  • Bolin Chen*
  • , Shanzhi Yin*
  • , Hanwei Zhu*
  • , Lingyu Zhu*
  • , Zihan Zhang*
  • , Jie Chen
  • , Ru Ling Liao
  • , Shiqi Wang*
  • , Yan Ye
  • *Corresponding author for this work
  • City University of Hong Kong
  • Alibaba Group Holding Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body signal into a series of configurable embeddings, which are controllably edited, compactly compressed, and efficiently transmitted. Moreover, the proposed decoder can evolve the mesh-based motion fields from these decoded semantics to realize the high-quality human body video reconstruction. Experimental results illustrate that the proposed framework can achieve promising compression performance for human body videos at ultra-low bitrate ranges compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes. Furthermore, the proposed framework enables interactive human body video coding without any additional pre-/post-manipulation processes, which is expected to shed light on metaverse-related digital human communication in the future.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PublisherIEEE Computer Society
Pages157-162
Number of pages6
ISBN (Electronic)9798331523794
DOIs
StatePublished - 2025
Externally publishedYes
Event32nd IEEE International Conference on Image Processing, ICIP 2025 - Anchorage, United States
Duration: 14 Sep 202517 Sep 2025

Publication series

NameProceedings - International Conference on Image Processing, ICIP
ISSN (Print)1522-4880

Conference

Conference32nd IEEE International Conference on Image Processing, ICIP 2025
Country/TerritoryUnited States
CityAnchorage
Period14/09/2517/09/25

Keywords

  • controllable embeddings
  • deep generative model
  • human body video
  • Interactive video coding

Fingerprint

Dive into the research topics of 'COMPRESSING HUMAN BODY VIDEO WITH INTERACTIVE SEMANTICS: A GENERATIVE APPROACH'. Together they form a unique fingerprint.

Cite this