TY - GEN
T1 - Reveal Fluidity Behind Frames
T2 - 26th IEEE International Workshop on Multimedia Signal Processing, MMSP 2024
AU - Xu, Siyuan
AU - Chen, Peilin
AU - Liu, Yue
AU - Wang, Meng
AU - Wang, Shiqi
AU - Kwong, Sam
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Assessing the quality of a player's performance, such as in diving events, requires precise measurement of subtle action details and overall fluidity. Existing methods primarily utilize appearance information from RGB frames, often neglecting crucial motion information that could contribute to a more comprehensive assessment. In response to this limitation, this paper introduces a novel Multi-Modality Network for Action Quality Assessment (AQA). The proposed method first employs a self-attention based module to foster interaction between optical flow and appearance clues, facilitating the extraction of discriminative features from each modality. Subsequently, a pairwise cross-attention mechanism is designed to comprehensively capture subtle differences via both intra-modality and inter-modality relationships between the query and exemplar video. Finally, to enhance the robustness and achieve accurate score prediction, an adaptive clip aggregation module is introduced to weigh the reliability of each patch based on multi-modal difference features. Experimental results on two benchmarks, FineDiving and MTL-AQA, validate the effectiveness of the proposed model.
AB - Assessing the quality of a player's performance, such as in diving events, requires precise measurement of subtle action details and overall fluidity. Existing methods primarily utilize appearance information from RGB frames, often neglecting crucial motion information that could contribute to a more comprehensive assessment. In response to this limitation, this paper introduces a novel Multi-Modality Network for Action Quality Assessment (AQA). The proposed method first employs a self-attention based module to foster interaction between optical flow and appearance clues, facilitating the extraction of discriminative features from each modality. Subsequently, a pairwise cross-attention mechanism is designed to comprehensively capture subtle differences via both intra-modality and inter-modality relationships between the query and exemplar video. Finally, to enhance the robustness and achieve accurate score prediction, an adaptive clip aggregation module is introduced to weigh the reliability of each patch based on multi-modal difference features. Experimental results on two benchmarks, FineDiving and MTL-AQA, validate the effectiveness of the proposed model.
KW - Action quality assessment
KW - Attention mechanism
KW - Multi-modal learning
UR - https://www.scopus.com/pages/publications/85211316404
U2 - 10.1109/MMSP61759.2024.10743311
DO - 10.1109/MMSP61759.2024.10743311
M3 - 会议稿件
AN - SCOPUS:85211316404
T3 - 2024 IEEE 26th International Workshop on Multimedia Signal Processing, MMSP 2024
BT - 2024 IEEE 26th International Workshop on Multimedia Signal Processing, MMSP 2024
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 2 October 2024 through 4 October 2024
ER -