跳到主要导航 跳到搜索 跳到主要内容

Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment

  • Baoliang Chen
  • , Lingyu Zhu
  • , Guo Li
  • , Fangbo Lu
  • , Hongfei Fan
  • , Shiqi Wang*
  • *此作品的通讯作者
  • City University of Hong Kong
  • Hongfei Fan Are with Kingsoft Cloud

科研成果: 期刊稿件文章同行评审

摘要

In this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method. The codes are released at https://github.com/Baoliang93/GSTVQA

源语言英语
页(从-至)1903-1916
页数14
期刊IEEE Transactions on Circuits and Systems for Video Technology
32
4
DOI
出版状态已出版 - 1 4月 2022
已对外发布

指纹

探究 'Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment' 的科研主题。它们共同构成独一无二的指纹。

引用此