TY - JOUR
T1 - Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment
AU - Chen, Baoliang
AU - Zhu, Lingyu
AU - Li, Guo
AU - Lu, Fangbo
AU - Fan, Hongfei
AU - Wang, Shiqi
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2022/4/1
Y1 - 2022/4/1
N2 - In this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method. The codes are released at https://github.com/Baoliang93/GSTVQA
AB - In this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method. The codes are released at https://github.com/Baoliang93/GSTVQA
KW - deep neural networks
KW - generalization capability
KW - temporal aggregation
KW - Video quality assessment
UR - https://www.scopus.com/pages/publications/85128630605
U2 - 10.1109/TCSVT.2021.3088505
DO - 10.1109/TCSVT.2021.3088505
M3 - 文章
AN - SCOPUS:85128630605
SN - 1051-8215
VL - 32
SP - 1903
EP - 1916
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
IS - 4
ER -