Skip to main navigation Skip to search Skip to main content

Towards Open-Ended Visual Quality Comparison

  • Haoning Wu*
  • , Hanwei Zhu
  • , Zicheng Zhang
  • , Erli Zhang
  • , Chaofeng Chen
  • , Liang Liao
  • , Chunyi Li
  • , Annan Wang
  • , Wenxiu Sun
  • , Qiong Yan
  • , Xiaohong Liu
  • , Guangtao Zhai
  • , Shiqi Wang
  • , Weisi Lin
  • *Corresponding author for this work
  • Nanyang Technological University
  • City University of Hong Kong
  • Shanghai Jiao Tong University
  • SenseTime Group Limited

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Comparative settings (e.g. pairwise choice, listwise ranking) have been adopted by a wide range of subjective studies for image quality assessment (IQA), as it inherently standardizes the evaluation criteria across different observers and offer more clear-cut responses. In this work, we extend the edge of emerging large multi-modality models (LMMs) to further advance visual quality comparison into open-ended settings, that 1) can respond to on quality comparison; 2) can provide beyond direct answers. To this end, we propose the Co-Instruct. To train this first-of-its-kind open-source open-ended visual quality comparer, we collect the Co-Instruct-562K dataset, from two sources: (a) LLM-merged single image quality description, (b) GPT-4V “teacher” responses on unlabeled data. Furthermore, to better evaluate this setting, we propose the MICBench, the first benchmark on multi-image comparison for LMMs. We demonstrate that Co-Instruct not only achieves in average 30% higher accuracy than state-of-the-art open-source LMMs, but also outperforms GPT-4V (its teacher), on both existing related benchmarks and the proposed MICBench. Our code, model and data are released on https://github.com/Q-Future/Co-Instruct.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages360-377
Number of pages18
ISBN (Print)9783031726453
DOIs
StatePublished - 2025
Externally publishedYes
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: 29 Sep 20244 Oct 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15061 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period29/09/244/10/24

Keywords

  • Large Multi-modality Models (LMM)
  • Visual Quality Assessment
  • Visual Quality Comparison
  • Visual Question Answering

Fingerprint

Dive into the research topics of 'Towards Open-Ended Visual Quality Comparison'. Together they form a unique fingerprint.

Cite this