跳到主要导航 跳到搜索 跳到主要内容

Model Merging for Knowledge Editing

  • Zichuan Fu
  • , Xian Wu*
  • , Guojing Li
  • , Yingying Zhang
  • , Yefeng Zheng
  • , Tianshi Ming
  • , Yejing Wang
  • , Wanyu Wang
  • , Xiangyu Zhao*
  • *此作品的通讯作者
  • City University of Hong Kong
  • Tencent
  • Westlake University
  • Tongji University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large Language Models (LLMs) require continuous updates to maintain accurate and current knowledge as the world evolves. While existing knowledge editing approaches offer various solutions for knowledge updating, they often struggle with sequential editing scenarios and harm the general capabilities of the model, thereby significantly hampering their practical applicability. This paper proposes a two-stage framework combining robust supervised fine-tuning (R-SFT) with model merging for knowledge editing. Our method first fine-tunes the LLM to internalize new knowledge fully, then merges the fine-tuned model with the original foundation model to preserve newly acquired knowledge and general capabilities. Experimental results demonstrate that our approach significantly outperforms existing methods in sequential editing while better preserving the original performance of the model, all without requiring any architectural changes. Code is available at Applied-Machine-Learning-Lab/MM4KE.

源语言英语
主期刊名Industry Track
编辑Georg Rehm, Yunyao Li
出版商Association for Computational Linguistics (ACL)
433-443
页数11
ISBN(电子版)9798891762886
DOI
出版状态已出版 - 2025
已对外发布
活动63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, 奥地利
期限: 27 7月 20251 8月 2025

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
6
ISSN(印刷版)0736-587X

会议

会议63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
国家/地区奥地利
Vienna
时期27/07/251/08/25

学术指纹

探究 'Model Merging for Knowledge Editing' 的科研主题。它们共同构成独一无二的学术指纹。

引用此