Skip to main navigation Skip to search Skip to main content

E2M-Fusion: An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data

  • Yuqiao Zeng
  • , Tengfei Liang
  • , Shiqi Wang
  • , Yidong Li
  • , Yi Jin*
  • , Qiuqi Ruan
  • *Corresponding author for this work
  • Beijing Jiaotong University
  • Ministry of Education of the People's Republic of China
  • City University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

Image fusion combines images from multiple sources to produce detailed results, but real-world situations often include unpaired samples with missing data due to issues like data privacy or camera differences. This challenge underscores the need to reconstruct unpaired samples and achieve effective fusion, improving decision-making in applications where data is incomplete. In this paper, we propose E2M-Fusion, the first End-to-end Multi-level model for multimodal image Fusion with missing data, which can bring unpaired sample reconstruction and image fusion into an end-to-end framework. E2M-Fusion operates at three levels: (1) a Supported Images Retrieval (SIR) strategy at the data level, leveraging paired samples and a vision-language model to retrieve reliable references for reconstruction; (2) a Fourier Contrastive Diffusion (FCD) model at the representation level, combining Fourier Transform, contrastive learning, and diffusion techniques for high-quality unpaired sample reconstruction; and (3) an Adaptive Labeling Fusion (ALF) model at the decision level, utilizing multi-task evaluation and iterative optimization to generate robust labels for fusion. Unlike existing methods limited to paired data, E2M-Fusion can address modality missing and fuses images in a unified framework. Extensive experiments on public datasets (M3FD, TNO, and LLVIP) demonstrate E2M Fusion can effectively handle the task of fusing unpaired images in real-world scenarios.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Denoising diffusion model
  • Fourier Transform
  • Image fusion
  • Missing modality reconstruction

Fingerprint

Dive into the research topics of 'E2M-Fusion: An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data'. Together they form a unique fingerprint.

Cite this