TY - GEN
T1 - How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?
AU - Fang, Xiaohan
AU - Chen, Peilin
AU - Wang, Meng
AU - Wang, Shiqi
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - The rapid evolution of Large Language Models (LLMs) has spurred the development of Large Vision Language Models (LVLMs), which demonstrate remarkable proficiency in various computer vision tasks through corresponding input prompts. These models perform impressively on diverse multimodal benchmarks, matching the effectiveness of conventional task-specific models. Nonetheless, many experiments with LVLMs often assume that the images are pristine, overlooking potential information loss during image transmission. To explore how image compression affects the semantic analysis capabilities of LVLMs in real-world scenarios, this paper presents a new image-text dataset named GPT-COMP. This dataset comprises 80,000 natural scene images, including 20,000 raw images from two public datasets. Each image is subjected to three different levels of compression distortion (QP = 32, 42, 52) using the latest Versatile Video Coding (VVC) Test Model. We leverage these variably compressed images in GPT-COMP to evaluate the state-of-the-art LVLM, GPT-4o, in vision understanding tasks. The text responses generated by GPT-4o are further organized into the GPT-COMP dataset and serve as the basis for evaluation. Specifically, the scene understanding capabilities regarding different compression levels are measured based on extracted semantic features and our proposed self-scoring strategy. This analysis sheds light on how image compression affects the semantic analysis capability of LLMs, offering valuable insights into the resilience of these models under realistic, suboptimal conditions.
AB - The rapid evolution of Large Language Models (LLMs) has spurred the development of Large Vision Language Models (LVLMs), which demonstrate remarkable proficiency in various computer vision tasks through corresponding input prompts. These models perform impressively on diverse multimodal benchmarks, matching the effectiveness of conventional task-specific models. Nonetheless, many experiments with LVLMs often assume that the images are pristine, overlooking potential information loss during image transmission. To explore how image compression affects the semantic analysis capabilities of LVLMs in real-world scenarios, this paper presents a new image-text dataset named GPT-COMP. This dataset comprises 80,000 natural scene images, including 20,000 raw images from two public datasets. Each image is subjected to three different levels of compression distortion (QP = 32, 42, 52) using the latest Versatile Video Coding (VVC) Test Model. We leverage these variably compressed images in GPT-COMP to evaluate the state-of-the-art LVLM, GPT-4o, in vision understanding tasks. The text responses generated by GPT-4o are further organized into the GPT-COMP dataset and serve as the basis for evaluation. Specifically, the scene understanding capabilities regarding different compression levels are measured based on extracted semantic features and our proposed self-scoring strategy. This analysis sheds light on how image compression affects the semantic analysis capability of LLMs, offering valuable insights into the resilience of these models under realistic, suboptimal conditions.
KW - compression distortion
KW - GPT-4o
KW - Large vision language models
KW - semantic evaluation
UR - https://www.scopus.com/pages/publications/85218188542
U2 - 10.1109/APSIPAASC63619.2025.10849295
DO - 10.1109/APSIPAASC63619.2025.10849295
M3 - 会议稿件
AN - SCOPUS:85218188542
T3 - APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024
BT - APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024
Y2 - 3 December 2024 through 6 December 2024
ER -