Skip to main navigation Skip to search Skip to main content

Function Calling in Large Language Models: Industrial Practices, Challenges, and Future Directions

  • Maolin Wang
  • , Yingyi Zhang
  • , Bowen Yu
  • , Bingguang Hao
  • , Cunyin Peng
  • , Yicheng Chen
  • , Wei Zhou
  • , Jinjie Gu
  • , Chenyi Zhuang
  • , Ruocheng Guo
  • , Wanyu Wang
  • , Xiangyu Zhao*
  • *Corresponding author for this work
  • City University of Hong Kong
  • Ant Group
  • Unaffliated

Research output: Contribution to journalArticlepeer-review

Abstract

The swift evolution of Large Language Models (LLMs) like the GPT family, LLaMA, ChatGLM, and Qwen represents significant progress in artificial intelligence research. Despite their remarkable capabilities in generating content, these models encounter substantial challenges when producing structured outputs and engaging in dynamic interactions, particularly when they need to retrieve external information in real time. To address these limitations, researchers have developed the “Function Calling” paradigm. This approach enables language models to analyze user inquiries and engage with defined functions, thereby facilitating precise responses through connections to external sources, including databases, programming interfaces, and live data streams. This functionality has been successfully implemented across numerous sectors such as finance analytics, healthcare systems, and service operations. The implementation of function calling comprises three essential phases: preparation, execution, and processing. The preparation phase encompasses query analysis and function identification. During execution, the system evaluates whether a function is necessary, extracts relevant parameters, and oversees the operation. The processing phase concentrates on analyzing outcomes and crafting appropriate responses. Each phase presents unique difficulties, ranging from accurately selecting functions to managing complex parameter extraction and ensuring reliable execution. Researchers have established various evaluation frameworks and metrics to assess function calling performance, including success rates, computational efficiency, parameter extraction accuracy, and response quality indicators such as ROUGE-L evaluation scores. This survey systematically reviews the current landscape of function calling in LLMs, analyzing technical challenges, examining existing solutions, and discussing evaluation methodologies. We particularly focus on practical implementations and industrial applications, providing insights into both current achievements and future directions in this rapidly evolving field. For a comprehensive collection of related research papers and the Appendix file, please refer to our repository at GitHub.<ani:xref

Original languageEnglish
Article number239
JournalACM Computing Surveys
Volume58
Issue number9
DOIs
StatePublished - Jul 2026
Externally publishedYes

Keywords

  • LLM agent
  • Large language models
  • function calling
  • industrial perspective

Fingerprint

Dive into the research topics of 'Function Calling in Large Language Models: Industrial Practices, Challenges, and Future Directions'. Together they form a unique fingerprint.

Cite this