Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Feb 16, 2024

Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, Xuanjing Huang

Figure 1 for ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Figure 2 for ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Figure 3 for ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Figure 4 for ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Share this with someone who'll enjoy it:

Abstract:Tool learning is widely acknowledged as a foundational approach or deploying large language models (LLMs) in real-world scenarios. While current research primarily emphasizes leveraging tools to augment LLMs, it frequently neglects emerging safety considerations tied to their application. To fill this gap, we present $ToolSword$, a comprehensive framework dedicated to meticulously investigating safety issues linked to LLMs in tool learning. Specifically, ToolSword delineates six safety scenarios for LLMs in tool learning, encompassing $malicious$ $queries$ and $jailbreak$ $attacks$ in the input stage, $noisy$ $misdirection$ and $risky$ $cues$ in the execution stage, and $harmful$ $feedback$ and $error$ $conflicts$ in the output stage. Experiments conducted on 11 open-source and closed-source LLMs reveal enduring safety challenges in tool learning, such as handling harmful queries, employing risky tools, and delivering detrimental feedback, which even GPT-4 is susceptible to. Moreover, we conduct further studies with the aim of fostering research on tool learning safety. The data is released in https://github.com/Junjie-Ye/ToolSword.

View paper on

Share this with someone who'll enjoy it:

Title:ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

Paper and Code