Skip to main navigation Skip to search Skip to main content

PAnDA: Combating Negative Augmentation via Large Language Models for User Cold-Start Recommendations

  • Yantong Du
  • , Rui Chen*
  • , Xiangyu Zhao*
  • , Qilong Han
  • , A. K. Qin
  • *Corresponding author for this work
  • Harbin Engineering University
  • City University of Hong Kong
  • Swinburne University of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The cold-start problem remains a long-standing challenge in recommender systems. Recent advances in large language models (LLMs) have opened new avenues for addressing cold-start scenarios through data augmentation. However, existing cold-start augmentation methods often suffer from negative augmentation, manifesting as incomplete augmentation, where generated interactions fail to comprehensively reflect user preferences, and inaccurate augmentation, where they conflict with user intent. These issues largely stem from two limitations: (1) the inability to effectively incorporate collaborative signals, which are critical for preference alignment, and (2) the lack of awareness of the downstream model's learning dynamics during data augmentation. To the best of our knowledge, the latter has not been studied in the literature. Consequently, we propose a novel framework named PAnDA. To address the incomplete augmentation issue, we propose a model-agnostic preference-aligned augmentation module to iteratively extract and fuse textual information and collaborative information by user-user preference matching and user-item preference coherence, which together form a contextual cue to guide the augmentor to generate high-quality augmented data. To overcome the inaccurate augmentation issue, we propose a model-specific downstream-model-aware adaptation module to adaptively align the augmented data with the model's states during the training process, guided by gradient similarity. Extensive experiments on three public benchmark datasets demonstrate that PAnDA outperforms different groups of state-of-the-art cold-start recommendation methods in all scenarios. The source code is publicly available at https://github.com/YantongDU/PAnDA.

Original languageEnglish
Title of host publicationCIKM 2025 - Proceedings of the 34th ACM International Conference on Information and Knowledge Management
PublisherAssociation for Computing Machinery, Inc
Pages3844-3854
Number of pages11
ISBN (Electronic)9798400720406
DOIs
StatePublished - 10 Nov 2025
Externally publishedYes
Event34th ACM International Conference on Information and Knowledge Management, CIKM 2025 - Seoul, Korea, Republic of
Duration: 10 Nov 202514 Nov 2025

Publication series

NameCIKM 2025 - Proceedings of the 34th ACM International Conference on Information and Knowledge Management

Conference

Conference34th ACM International Conference on Information and Knowledge Management, CIKM 2025
Country/TerritoryKorea, Republic of
CitySeoul
Period10/11/2514/11/25

Keywords

  • cold-start recommendations
  • data augmentation
  • large language models
  • meta-learning

Fingerprint

Dive into the research topics of 'PAnDA: Combating Negative Augmentation via Large Language Models for User Cold-Start Recommendations'. Together they form a unique fingerprint.

Cite this