Skip to main navigation Skip to search Skip to main content

How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?

  • City University of Hong Kong

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The rapid evolution of Large Language Models (LLMs) has spurred the development of Large Vision Language Models (LVLMs), which demonstrate remarkable proficiency in various computer vision tasks through corresponding input prompts. These models perform impressively on diverse multimodal benchmarks, matching the effectiveness of conventional task-specific models. Nonetheless, many experiments with LVLMs often assume that the images are pristine, overlooking potential information loss during image transmission. To explore how image compression affects the semantic analysis capabilities of LVLMs in real-world scenarios, this paper presents a new image-text dataset named GPT-COMP. This dataset comprises 80,000 natural scene images, including 20,000 raw images from two public datasets. Each image is subjected to three different levels of compression distortion (QP = 32, 42, 52) using the latest Versatile Video Coding (VVC) Test Model. We leverage these variably compressed images in GPT-COMP to evaluate the state-of-the-art LVLM, GPT-4o, in vision understanding tasks. The text responses generated by GPT-4o are further organized into the GPT-COMP dataset and serve as the basis for evaluation. Specifically, the scene understanding capabilities regarding different compression levels are measured based on extracted semantic features and our proposed self-scoring strategy. This analysis sheds light on how image compression affects the semantic analysis capability of LLMs, offering valuable insights into the resilience of these models under realistic, suboptimal conditions.

Original languageEnglish
Title of host publicationAPSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798350367331
DOIs
StatePublished - 2024
Externally publishedYes
Event2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024 - Macau, China
Duration: 3 Dec 20246 Dec 2024

Publication series

NameAPSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024

Conference

Conference2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024
Country/TerritoryChina
CityMacau
Period3/12/246/12/24

Keywords

  • compression distortion
  • GPT-4o
  • Large vision language models
  • semantic evaluation

Fingerprint

Dive into the research topics of 'How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?'. Together they form a unique fingerprint.

Cite this