跳到主要导航 跳到搜索 跳到主要内容

Stepwise Reasoning Disruption Attack of LLMs

  • Jingyu Peng
  • , Maolin Wang
  • , Xiangyu Zhao*
  • , Kai Zhang
  • , Wanyu Wang
  • , Pengyue Jia
  • , Qidong Liu
  • , Ruocheng Guo
  • , Qi Liu*
  • *此作品的通讯作者
  • University of Science and Technology of China
  • City University of Hong Kong

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large language models (LLMs) have made remarkable strides in complex reasoning tasks, but their safety and robustness in reasoning processes remain unexplored, particularly in third-party platforms that facilitate user interactions via APIs. Existing attacks on LLM reasoning are constrained by specific settings or lack of imperceptibility, limiting their feasibility and generalizability. To address these challenges, we propose the Stepwise rEasoning Error Disruption (SEED) attack, which subtly injects errors into prior reasoning steps to mislead the model into producing incorrect subsequent reasoning and final answers. Unlike previous methods, SEED is compatible with zero-shot and few-shot settings, maintains the natural reasoning flow, and ensures covert execution without modifying the instruction. Extensive experiments on four datasets across four different models demonstrate SEED's effectiveness, revealing the vulnerabilities of LLMs to disruptions in reasoning processes. These findings underscore the need for greater attention to the robustness of LLM reasoning to ensure safety in practical applications. Our code is available at: https://github.com/Applied-Machine-Learning-Lab/SEED-Attack.

源语言英语
主期刊名Long Papers
编辑Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
出版商Association for Computational Linguistics (ACL)
5040-5058
页数19
ISBN(电子版)9798891762510
DOI
出版状态已出版 - 2025
已对外发布
活动63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, 奥地利
期限: 27 7月 20251 8月 2025

丛书

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
1
ISSN(印刷版)0736-587X

会议

会议63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
国家/地区奥地利
Vienna
时期27/07/251/08/25

学术指纹

探究 'Stepwise Reasoning Disruption Attack of LLMs' 的科研主题。它们共同构成独一无二的学术指纹。

引用此