Skip to main navigation Skip to search Skip to main content

NoteLLM-2: Multimodal Large Representation Models for Recommendation

  • Chao Zhang
  • , Haoxin Zhang
  • , Shiwei Wu
  • , Di Wu
  • , Tong Xu*
  • , Xiangyu Zhao*
  • , Yan Gao
  • , Yao Hu
  • , Enhong Chen*
  • *Corresponding author for this work
  • University of Science and Technology of China
  • Xiaohongshu
  • City University of Hong Kong

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularly for item-to-item (I2I) recommendations, remains underexplored. While leveraging existing Multimodal Large Language Models (MLLMs) for such tasks is promising, challenges arise due to their delayed release compared to corresponding LLMs and the inefficiency in representation tasks. To address these issues, we propose an end-to-end fine-tuning method that customizes the integration of any existing LLMs and vision encoders for efficient multimodal representation. Preliminary experiments revealed that fine-tuned LLMs often neglect image content. To counteract this, we propose NoteLLM-2, a novel framework that enhances visual information. Specifically, we propose two approaches: first, a prompt-based method that segregates visual and textual content, employing a multimodal In-Context Learning strategy to balance focus across modalities; second, a late fusion technique that directly integrates visual information into the final representations. Extensive experiments, both online and offline, demonstrate the effectiveness of our approach. Code is available at https://github.com/Applied-Machine-Learning-Lab/NoteLLM.

Original languageEnglish
Title of host publicationKDD 2025 - Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
PublisherAssociation for Computing Machinery
Pages2815-2826
Number of pages12
ISBN (Electronic)9798400712456
DOIs
StatePublished - 20 Jul 2025
Externally publishedYes
Event31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025 - Toronto, Canada
Duration: 3 Aug 20257 Aug 2025

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume1
ISSN (Print)2154-817X

Conference

Conference31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025
Country/TerritoryCanada
CityToronto
Period3/08/257/08/25

Keywords

  • multimodal large language model
  • multimodal representation
  • recommendation

Fingerprint

Dive into the research topics of 'NoteLLM-2: Multimodal Large Representation Models for Recommendation'. Together they form a unique fingerprint.

Cite this