Skip to main navigation Skip to search Skip to main content

Multi-modal Recommendation with Joint Content and Interaction Augmentation

  • Jiajie Deng
  • , Haokun Wen
  • , Xiao Han
  • , Xuemeng Song
  • , Xiangyu Zhao*
  • *Corresponding author for this work
  • City University of Hong Kong
  • Harbin Institute of Technology Shenzhen
  • Zhejiang University of Technology
  • Southern University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Multi-modal recommender systems have become indispensable in modern applications. Despite promising results, existing methods have two key limitations. First, in item characteristic modeling, they rely solely on the merchant's description and often neglect customer reviews, leading to biased item quality assessment. Second, in user preference modeling, they focus mainly on one-hop user-item interactions in the interaction graph, overlooking multi-hop interactions, which limits the understanding of user preferences. To address these issues, we propose a Joint Content and Interaction Augmented Framework (JCIAF) for multi-modal recommendation. Specifically, we leverage large language models to extract valuable insights from user reviews, integrating this with the merchant's description to form a more comprehensive textual representation of the item. This enriched description provides a balanced foundation for item characteristic modeling. Next, we enhance the user-item interaction graph with two additional interaction types: user-item-item (two-hop) and user-item-user-item (three-hop), which offer augmented views for more thorough user preference modeling. We apply a diffusion-based method across the three augmented graphs and introduce an online knowledge distillation mechanism to enable cross-graph learning. Extensive experiments on three real-world datasets demonstrate the effectiveness of our proposed method. The source code is accessible at https://github.com/jjlinnn/JClAF.git.

Original languageEnglish
Title of host publicationProceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
EditorsTat-Seng Chua, Lai-Kuan Wong, Chee Seng Chan, Jinhui Tang, Chong-Wah Ngo, Klaus Schoeffmann, Jiaying Liu, Yo-Sung Ho
PublisherAssociation for Computing Machinery, Inc
ISBN (Electronic)9798400720055
DOIs
StatePublished - 6 Dec 2025
Externally publishedYes
Event7th ACM International Conference on Multimedia in Asia, MMAsia 2025 - Kuala Lumpur, Malaysia
Duration: 9 Dec 202512 Dec 2025

Publication series

NameProceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025

Conference

Conference7th ACM International Conference on Multimedia in Asia, MMAsia 2025
Country/TerritoryMalaysia
CityKuala Lumpur
Period9/12/2512/12/25

Keywords

  • Multi-modal Recommender Systems
  • Multi-perspective Modeling
  • Mutual Enhancement

Fingerprint

Dive into the research topics of 'Multi-modal Recommendation with Joint Content and Interaction Augmentation'. Together they form a unique fingerprint.

Cite this