TY - GEN
T1 - Disentangled human action video generation via decoupled learning
AU - Yang, Lingbo
AU - Zhao, Zhenghui
AU - Wang, Shiqi
AU - Wang, Shanshe
AU - Ma, Siwei
AU - Gao, Wen
N1 - Publisher Copyright:
© 2019 IEEE.
PY - 2019/7
Y1 - 2019/7
N2 - Recently there has been remarkable progress in synthesizing realistic human action videos by directly learning to translate pose heatmaps/stick figures to video frames in an end-to-end fashion. However, such models are not suitable for fashion-related applications that typically require flexible manipulations of visual attributes, such as the color of clothes. In this paper, we propose a disentangled human video generation framework conditioned on both the pose sequence and encoded color attributes. We aim to learn an encoder that captures the manifold structure of latent color space and a generator that fully utilizes the encoded color attributes to produce diversely-colored human action videos. To this end, we design a two-stage decoupled learning approach that uses a pre-trained color-aware encoder to guide the disentangled learning of the generator. Furthermore, a color augmentation approach is applied on raw video clips to better shape the distribution of samples in the latent color space. Comprehensive experimental results demonstrate the efficacy of our proposed methods.
AB - Recently there has been remarkable progress in synthesizing realistic human action videos by directly learning to translate pose heatmaps/stick figures to video frames in an end-to-end fashion. However, such models are not suitable for fashion-related applications that typically require flexible manipulations of visual attributes, such as the color of clothes. In this paper, we propose a disentangled human video generation framework conditioned on both the pose sequence and encoded color attributes. We aim to learn an encoder that captures the manifold structure of latent color space and a generator that fully utilizes the encoded color attributes to produce diversely-colored human action videos. To this end, we design a two-stage decoupled learning approach that uses a pre-trained color-aware encoder to guide the disentangled learning of the generator. Furthermore, a color augmentation approach is applied on raw video clips to better shape the distribution of samples in the latent color space. Comprehensive experimental results demonstrate the efficacy of our proposed methods.
KW - Decoupled learning
KW - Feature disentanglement
KW - Generative adversarial networks (GANs)
KW - Pose-guided video generation
UR - https://www.scopus.com/pages/publications/85071459423
U2 - 10.1109/ICMEW.2019.00091
DO - 10.1109/ICMEW.2019.00091
M3 - 会议稿件
AN - SCOPUS:85071459423
T3 - Proceedings - 2019 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2019
SP - 495
EP - 500
BT - Proceedings - 2019 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2019
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2019 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2019
Y2 - 8 July 2019 through 12 July 2019
ER -