TY - GEN
T1 - The Elephant in the Room
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Guo, Xinwei
AU - Gao, Jiashi
AU - Zhou, Junlei
AU - Zhang, Jiaxin
AU - Chen, Guanhua
AU - Zhao, Xiangyu
AU - Liu, Quanying
AU - Wu, Haiyan
AU - Yao, Xin
AU - Wei, Xuetao
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Large language models (LLMs) are increasingly integrated into our daily lives, raising significant ethical concerns, especially about perpetuating stereotypes. While group-specific debiasing methods have made progress, they often fail to address multiple biases simultaneously. In contrast, group-agnostic debiasing has the potential to mitigate a variety of biases at once, but remains underexplored. In this work, we investigate the role of neutral words-the group-agnostic component-in enhancing the group-agnostic debiasing process. We first reveal that neutral words are essential for preserving semantic modeling, and we propose ϵ-DPCE, a method that incorporates a neutral word semantics-based loss function to effectively alleviate the deterioration of the Language Modeling Score (LMS) during the debiasing process. Furthermore, by introducing the SCM-Projection method, we demonstrate that SCM-based debiasing eliminates stereotypes by indirectly disrupting the association between attribute and neutral words in the Stereotype Content Model (SCM) space. Our experiments show that neutral words, which often embed multi-group stereotypical objects, play a key role in contributing to the group-agnostic nature of SCM-based debiasing.
AB - Large language models (LLMs) are increasingly integrated into our daily lives, raising significant ethical concerns, especially about perpetuating stereotypes. While group-specific debiasing methods have made progress, they often fail to address multiple biases simultaneously. In contrast, group-agnostic debiasing has the potential to mitigate a variety of biases at once, but remains underexplored. In this work, we investigate the role of neutral words-the group-agnostic component-in enhancing the group-agnostic debiasing process. We first reveal that neutral words are essential for preserving semantic modeling, and we propose ϵ-DPCE, a method that incorporates a neutral word semantics-based loss function to effectively alleviate the deterioration of the Language Modeling Score (LMS) during the debiasing process. Furthermore, by introducing the SCM-Projection method, we demonstrate that SCM-based debiasing eliminates stereotypes by indirectly disrupting the association between attribute and neutral words in the Stereotype Content Model (SCM) space. Our experiments show that neutral words, which often embed multi-group stereotypical objects, play a key role in contributing to the group-agnostic nature of SCM-based debiasing.
UR - https://www.scopus.com/pages/publications/105028615689
U2 - 10.18653/v1/2025.findings-acl.1044
DO - 10.18653/v1/2025.findings-acl.1044
M3 - 会议稿件
AN - SCOPUS:105028615689
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 20360
EP - 20371
BT - Findings of the Association for Computational Linguistics
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -