跳到主要导航 跳到搜索 跳到主要内容

How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?

  • City University of Hong Kong

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The rapid evolution of Large Language Models (LLMs) has spurred the development of Large Vision Language Models (LVLMs), which demonstrate remarkable proficiency in various computer vision tasks through corresponding input prompts. These models perform impressively on diverse multimodal benchmarks, matching the effectiveness of conventional task-specific models. Nonetheless, many experiments with LVLMs often assume that the images are pristine, overlooking potential information loss during image transmission. To explore how image compression affects the semantic analysis capabilities of LVLMs in real-world scenarios, this paper presents a new image-text dataset named GPT-COMP. This dataset comprises 80,000 natural scene images, including 20,000 raw images from two public datasets. Each image is subjected to three different levels of compression distortion (QP = 32, 42, 52) using the latest Versatile Video Coding (VVC) Test Model. We leverage these variably compressed images in GPT-COMP to evaluate the state-of-the-art LVLM, GPT-4o, in vision understanding tasks. The text responses generated by GPT-4o are further organized into the GPT-COMP dataset and serve as the basis for evaluation. Specifically, the scene understanding capabilities regarding different compression levels are measured based on extracted semantic features and our proposed self-scoring strategy. This analysis sheds light on how image compression affects the semantic analysis capability of LLMs, offering valuable insights into the resilience of these models under realistic, suboptimal conditions.

源语言英语
主期刊名APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798350367331
DOI
出版状态已出版 - 2024
已对外发布
活动2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024 - Macau, 中国
期限: 3 12月 20246 12月 2024

出版系列

姓名APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024

会议

会议2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024
国家/地区中国
Macau
时期3/12/246/12/24

学术指纹

探究 'How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?' 的科研主题。它们共同构成独一无二的学术指纹。

引用此