Skip to main navigation Skip to search Skip to main content

2AFC Prompting of Large Multimodal Models for Image Quality Assessment

  • Hanwei Zhu
  • , Xiangjie Sui
  • , Baoliang Chen
  • , Xuelin Liu
  • , Peilin Chen
  • , Yuming Fang
  • , Shiqi Wang*
  • *Corresponding author for this work
  • City University of Hong Kong
  • City University of Macau
  • South China Normal University
  • Jiangxi University of Finance and Economics

Research output: Contribution to journalArticlepeer-review

Abstract

While abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models (LMMs), their image quality assessment (IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice (2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posteriori estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available at https://github.com/h4nwei/2AFC-LMMs.

Original languageEnglish
Pages (from-to)12873-12878
Number of pages6
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume34
Issue number12
DOIs
StatePublished - 2024
Externally publishedYes

Keywords

  • image quality assessment
  • Large multimodal models
  • two-alternative forced choice

Fingerprint

Dive into the research topics of '2AFC Prompting of Large Multimodal Models for Image Quality Assessment'. Together they form a unique fingerprint.

Cite this