TY - GEN
T1 - Exploration and Regularization of the Latent Action Space in Recommendation
AU - Liu, Shuchang
AU - Cai, Qingpeng
AU - Sun, Bowen
AU - Wang, Yuhao
AU - Jiang, Ji
AU - Zheng, Dong
AU - Jiang, Peng
AU - Gai, Kun
AU - Zhao, Xiangyu
AU - Zhang, Yongfeng
N1 - Publisher Copyright:
© 2023 ACM.
PY - 2023/4/30
Y1 - 2023/4/30
N2 - In recommender systems, reinforcement learning solutions have effectively boosted recommendation performance because of their ability to capture long-term user-system interaction. However, the action space of the recommendation policy is a list of items, which could be extremely large with a dynamic candidate item pool. To overcome this challenge, we propose a hyper-actor and critic learning framework where the policy decomposes the item list generation process into a hyper-action inference step and an effect-action selection step. The first step maps the given state space into a vectorized hyper-action space, and the second step selects the item list based on the hyper-action. In order to regulate the discrepancy between the two action spaces, we design an alignment module along with a kernel mapping function for items to ensure inference accuracy and include a supervision module to stabilize the learning process. We build simulated environments on public datasets and empirically show that our framework is superior in recommendation compared to standard RL baselines.
AB - In recommender systems, reinforcement learning solutions have effectively boosted recommendation performance because of their ability to capture long-term user-system interaction. However, the action space of the recommendation policy is a list of items, which could be extremely large with a dynamic candidate item pool. To overcome this challenge, we propose a hyper-actor and critic learning framework where the policy decomposes the item list generation process into a hyper-action inference step and an effect-action selection step. The first step maps the given state space into a vectorized hyper-action space, and the second step selects the item list based on the hyper-action. In order to regulate the discrepancy between the two action spaces, we design an alignment module along with a kernel mapping function for items to ensure inference accuracy and include a supervision module to stabilize the learning process. We build simulated environments on public datasets and empirically show that our framework is superior in recommendation compared to standard RL baselines.
KW - Recommender Systems
KW - Reinforcement Learning
KW - Representation Learning
UR - https://www.scopus.com/pages/publications/85159300358
U2 - 10.1145/3543507.3583244
DO - 10.1145/3543507.3583244
M3 - 会议稿件
AN - SCOPUS:85159300358
T3 - ACM Web Conference 2023 - Proceedings of the World Wide Web Conference, WWW 2023
SP - 833
EP - 844
BT - ACM Web Conference 2023 - Proceedings of the World Wide Web Conference, WWW 2023
PB - Association for Computing Machinery, Inc
T2 - 32nd ACM World Wide Web Conference, WWW 2023
Y2 - 30 April 2023 through 4 May 2023
ER -