TY - JOUR
T1 - Monotonic and Invertible Network
T2 - A General Framework for Learning IQA Model from Mixed Datasets
AU - Chen, Baoliang
AU - Xiao, Kang
AU - Shen, Xuelin
AU - Wang, Shiqi
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2025.
PY - 2025/11
Y1 - 2025/11
N2 - Learning from mixed datasets is a powerful strategy for model generalization improvement. However, this approach becomes particularly challenging for image quality assessment (IQA) where the quality annotations across different datasets are usually not aligned, due to varying quality criteria, score ranges, and viewing conditions involved in their subjective tests. Score rescaling and rank learning are two main strategies for the mixed dataset IQA model learning. In particular, the score rescaling attempts to align the annotations across datasets directly by empirical linear or nonlinear transformations, which may not always be reliable. Rank learning, on the other hand, is restricted to pair comparison within each dataset while the image pairs from different datasets are not fully examined. In this paper, we present a novel mixed dataset learning framework for the IQA, where we align the quality annotation in an implicit manner. Specifically, we first introduce a dataset-shared quality regressor to project images across datasets into a unified proxy score space. We force the proxy scores to serve as proxies of the aligned results of the annotations. To achieve this, two key priors are considered: 1) within each dataset, the proxy scores should maintain the same rank as the annotations; and 2) the score ranges of proxy scores from different datasets overlap when images of similar quality exist in these datasets. To meet the criteria, we propose a monotonic and invertible network as a dataset-specific score mapper to bridge the proxy scores and the annotations with the rank consistency and range intersection established. This constrained learning ultimately results in the proxy scores being the desired alignment results of the annotations. Experiments on extensive IQA datasets validate the effectiveness of our method, and the continuous performance gains when incorporating our learning strategy into different network architectures demonstrate the high generalizability of our framework. The source code is available at https://github.com/KANGX99/MIMI.
AB - Learning from mixed datasets is a powerful strategy for model generalization improvement. However, this approach becomes particularly challenging for image quality assessment (IQA) where the quality annotations across different datasets are usually not aligned, due to varying quality criteria, score ranges, and viewing conditions involved in their subjective tests. Score rescaling and rank learning are two main strategies for the mixed dataset IQA model learning. In particular, the score rescaling attempts to align the annotations across datasets directly by empirical linear or nonlinear transformations, which may not always be reliable. Rank learning, on the other hand, is restricted to pair comparison within each dataset while the image pairs from different datasets are not fully examined. In this paper, we present a novel mixed dataset learning framework for the IQA, where we align the quality annotation in an implicit manner. Specifically, we first introduce a dataset-shared quality regressor to project images across datasets into a unified proxy score space. We force the proxy scores to serve as proxies of the aligned results of the annotations. To achieve this, two key priors are considered: 1) within each dataset, the proxy scores should maintain the same rank as the annotations; and 2) the score ranges of proxy scores from different datasets overlap when images of similar quality exist in these datasets. To meet the criteria, we propose a monotonic and invertible network as a dataset-specific score mapper to bridge the proxy scores and the annotations with the rank consistency and range intersection established. This constrained learning ultimately results in the proxy scores being the desired alignment results of the annotations. Experiments on extensive IQA datasets validate the effectiveness of our method, and the continuous performance gains when incorporating our learning strategy into different network architectures demonstrate the high generalizability of our framework. The source code is available at https://github.com/KANGX99/MIMI.
KW - Image quality assessment
KW - Invertible neural network
KW - Mixed dataset
KW - Monotonic network
UR - https://www.scopus.com/pages/publications/105013578482
U2 - 10.1007/s11263-025-02565-6
DO - 10.1007/s11263-025-02565-6
M3 - 文章
AN - SCOPUS:105013578482
SN - 0920-5691
VL - 133
SP - 7924
EP - 7945
JO - International Journal of Computer Vision
JF - International Journal of Computer Vision
IS - 11
ER -