跳到主要导航 跳到搜索 跳到主要内容

Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework

  • Xiao Han
  • , Chen Zhu*
  • , Hengshu Zhu*
  • , Xiangyu Zhao*
  • *此作品的通讯作者
  • City University of Hong Kong
  • University of Science and Technology of China
  • CAS - Computer Network Information Center

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Visual geo-localization demands in-depth knowledge and advanced reasoning skills to associate images with precise real-world geographic locations. Existing image database retrieval methods are limited by the impracticality of storing sufficient visual records of global landmarks. Recently, Large Vision-Language Models (LVLMs) have demonstrated the capability of geo-localization through Visual Question Answering (VQA), enabling a solution that does not require external geo-tagged image records. However, the performance of a single LVLM is still limited by its intrinsic knowledge and reasoning capabilities. To address these challenges, we introduce smileGeo, a novel visual geo-localization framework that leverages multiple Internet-enabled LVLM agents operating within an agent-based architecture. By facilitating inter-agent communication, smileGeo integrates the inherent knowledge of these agents with additional retrieved information, enhancing the ability to effectively localize images. Furthermore, our framework incorporates a dynamic learning strategy that optimizes agent communication, reducing redundant interactions and enhancing overall system efficiency. To validate the effectiveness of the proposed framework, we conducted experiments on three different datasets, and the results show that our approach significantly outperforms current state-of-the-art methods.

源语言英语
主期刊名KDD 2025 - Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
出版商Association for Computing Machinery
814-825
页数12
ISBN(电子版)9798400714542
DOI
出版状态已出版 - 3 8月 2025
已对外发布
活动31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025 - Toronto, 加拿大
期限: 3 8月 20257 8月 2025

出版系列

姓名Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
2
ISSN(印刷版)2154-817X

会议

会议31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025
国家/地区加拿大
Toronto
时期3/08/257/08/25

指纹

探究 'Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework' 的科研主题。它们共同构成独一无二的指纹。

引用此