TY - JOUR
T1 - E2M-Fusion
T2 - An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data
AU - Zeng, Yuqiao
AU - Liang, Tengfei
AU - Wang, Shiqi
AU - Li, Yidong
AU - Jin, Yi
AU - Ruan, Qiuqi
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Image fusion combines images from multiple sources to produce detailed results, but real-world situations often include unpaired samples with missing data due to issues like data privacy or camera differences. This challenge underscores the need to reconstruct unpaired samples and achieve effective fusion, improving decision-making in applications where data is incomplete. In this paper, we propose E2M-Fusion, the first End-to-end Multi-level model for multimodal image Fusion with missing data, which can bring unpaired sample reconstruction and image fusion into an end-to-end framework. E2M-Fusion operates at three levels: (1) a Supported Images Retrieval (SIR) strategy at the data level, leveraging paired samples and a vision-language model to retrieve reliable references for reconstruction; (2) a Fourier Contrastive Diffusion (FCD) model at the representation level, combining Fourier Transform, contrastive learning, and diffusion techniques for high-quality unpaired sample reconstruction; and (3) an Adaptive Labeling Fusion (ALF) model at the decision level, utilizing multi-task evaluation and iterative optimization to generate robust labels for fusion. Unlike existing methods limited to paired data, E2M-Fusion can address modality missing and fuses images in a unified framework. Extensive experiments on public datasets (M3FD, TNO, and LLVIP) demonstrate E2M Fusion can effectively handle the task of fusing unpaired images in real-world scenarios.
AB - Image fusion combines images from multiple sources to produce detailed results, but real-world situations often include unpaired samples with missing data due to issues like data privacy or camera differences. This challenge underscores the need to reconstruct unpaired samples and achieve effective fusion, improving decision-making in applications where data is incomplete. In this paper, we propose E2M-Fusion, the first End-to-end Multi-level model for multimodal image Fusion with missing data, which can bring unpaired sample reconstruction and image fusion into an end-to-end framework. E2M-Fusion operates at three levels: (1) a Supported Images Retrieval (SIR) strategy at the data level, leveraging paired samples and a vision-language model to retrieve reliable references for reconstruction; (2) a Fourier Contrastive Diffusion (FCD) model at the representation level, combining Fourier Transform, contrastive learning, and diffusion techniques for high-quality unpaired sample reconstruction; and (3) an Adaptive Labeling Fusion (ALF) model at the decision level, utilizing multi-task evaluation and iterative optimization to generate robust labels for fusion. Unlike existing methods limited to paired data, E2M-Fusion can address modality missing and fuses images in a unified framework. Extensive experiments on public datasets (M3FD, TNO, and LLVIP) demonstrate E2M Fusion can effectively handle the task of fusing unpaired images in real-world scenarios.
KW - Denoising diffusion model
KW - Fourier Transform
KW - Image fusion
KW - Missing modality reconstruction
UR - https://www.scopus.com/pages/publications/105041985825
U2 - 10.1109/TMM.2026.3701558
DO - 10.1109/TMM.2026.3701558
M3 - 文章
AN - SCOPUS:105041985825
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -