跳到主要导航 跳到搜索 跳到主要内容

E2M-Fusion: An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data

  • Yuqiao Zeng
  • , Tengfei Liang
  • , Shiqi Wang
  • , Yidong Li
  • , Yi Jin*
  • , Qiuqi Ruan
  • *此作品的通讯作者
  • Beijing Jiaotong University
  • Ministry of Education of the People's Republic of China
  • City University of Hong Kong

科研成果: 期刊稿件文章同行评审

摘要

Image fusion combines images from multiple sources to produce detailed results, but real-world situations often include unpaired samples with missing data due to issues like data privacy or camera differences. This challenge underscores the need to reconstruct unpaired samples and achieve effective fusion, improving decision-making in applications where data is incomplete. In this paper, we propose E2M-Fusion, the first End-to-end Multi-level model for multimodal image Fusion with missing data, which can bring unpaired sample reconstruction and image fusion into an end-to-end framework. E2M-Fusion operates at three levels: (1) a Supported Images Retrieval (SIR) strategy at the data level, leveraging paired samples and a vision-language model to retrieve reliable references for reconstruction; (2) a Fourier Contrastive Diffusion (FCD) model at the representation level, combining Fourier Transform, contrastive learning, and diffusion techniques for high-quality unpaired sample reconstruction; and (3) an Adaptive Labeling Fusion (ALF) model at the decision level, utilizing multi-task evaluation and iterative optimization to generate robust labels for fusion. Unlike existing methods limited to paired data, E2M-Fusion can address modality missing and fuses images in a unified framework. Extensive experiments on public datasets (M3FD, TNO, and LLVIP) demonstrate E2M Fusion can effectively handle the task of fusing unpaired images in real-world scenarios.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026
已对外发布

指纹

探究 'E2M-Fusion: An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data' 的科研主题。它们共同构成独一无二的指纹。

引用此