TY - JOUR
T1 - Enhanced motion-compensated video coding with deep virtual reference frame generation
AU - Wang, Shanshe
AU - Zhao, Lei
AU - Wang, Shiqi
AU - Zhang, Xinfeng
AU - Ma, Siwei
AU - Gao, Wen
N1 - Publisher Copyright:
© 1992-2012 IEEE.
PY - 2019/10
Y1 - 2019/10
N2 - In this paper, we propose an efficient inter prediction scheme by introducing the deep virtual reference frame (VRF), which serves better reference in the temporal redundancy removal process of video coding. In particular, the high quality VRF is generated with the deep learning-based frame rate up conversion (FRUC) algorithm from two reconstructed bi-directional frames, which is subsequently incorporated into the reference list serving as the high quality reference. Moreover, to alleviate the compression artifacts of VRF, we develop a convolutional neural network (CNN)-based enhancement model to further improve its quality. To facilitate better utilization of the VRF, a CTU level coding mode termed as direct virtual reference frame (DVRF) is devised, which achieves better trade-off between compression performance and complexity. The proposed scheme is integrated into HM-16.6 and JEM-7.1 software platforms, and the simulation results under random access (RA) configuration demonstrate significant superiority of the proposed method. When adding VRF to RPS, more than 6% average BD-rate gain is achieved for HEVC test sequences on HM-16.6, and 0.8% BD-rate gain is observed based on JEM-7.1 software. Regarding the DVRF mode, 3.6% bitrate saving is achieved on HM-16.6 with the computational complexity effectively reduced.
AB - In this paper, we propose an efficient inter prediction scheme by introducing the deep virtual reference frame (VRF), which serves better reference in the temporal redundancy removal process of video coding. In particular, the high quality VRF is generated with the deep learning-based frame rate up conversion (FRUC) algorithm from two reconstructed bi-directional frames, which is subsequently incorporated into the reference list serving as the high quality reference. Moreover, to alleviate the compression artifacts of VRF, we develop a convolutional neural network (CNN)-based enhancement model to further improve its quality. To facilitate better utilization of the VRF, a CTU level coding mode termed as direct virtual reference frame (DVRF) is devised, which achieves better trade-off between compression performance and complexity. The proposed scheme is integrated into HM-16.6 and JEM-7.1 software platforms, and the simulation results under random access (RA) configuration demonstrate significant superiority of the proposed method. When adding VRF to RPS, more than 6% average BD-rate gain is achieved for HEVC test sequences on HM-16.6, and 0.8% BD-rate gain is observed based on JEM-7.1 software. Regarding the DVRF mode, 3.6% bitrate saving is achieved on HM-16.6 with the computational complexity effectively reduced.
KW - deep learning
KW - Inter prediction
KW - video coding
KW - virtual reference frame
UR - https://www.scopus.com/pages/publications/85070439622
U2 - 10.1109/TIP.2019.2913545
DO - 10.1109/TIP.2019.2913545
M3 - 文章
C2 - 31059444
AN - SCOPUS:85070439622
SN - 1057-7149
VL - 28
SP - 4832
EP - 4844
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
IS - 10
M1 - 8704997
ER -