Abstract
Image fusion combines images from multiple sources to produce detailed results, but real-world situations often include unpaired samples with missing data due to issues like data privacy or camera differences. This challenge underscores the need to reconstruct unpaired samples and achieve effective fusion, improving decision-making in applications where data is incomplete. In this paper, we propose E2M-Fusion, the first End-to-end Multi-level model for multimodal image Fusion with missing data, which can bring unpaired sample reconstruction and image fusion into an end-to-end framework. E2M-Fusion operates at three levels: (1) a Supported Images Retrieval (SIR) strategy at the data level, leveraging paired samples and a vision-language model to retrieve reliable references for reconstruction; (2) a Fourier Contrastive Diffusion (FCD) model at the representation level, combining Fourier Transform, contrastive learning, and diffusion techniques for high-quality unpaired sample reconstruction; and (3) an Adaptive Labeling Fusion (ALF) model at the decision level, utilizing multi-task evaluation and iterative optimization to generate robust labels for fusion. Unlike existing methods limited to paired data, E2M-Fusion can address modality missing and fuses images in a unified framework. Extensive experiments on public datasets (M3FD, TNO, and LLVIP) demonstrate E2M Fusion can effectively handle the task of fusing unpaired images in real-world scenarios.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Denoising diffusion model
- Fourier Transform
- Image fusion
- Missing modality reconstruction
Fingerprint
Dive into the research topics of 'E2M-Fusion: An End-to-end Multi-level Model for Multimodal Image Fusion with Missing Data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver