Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Qinqyun Wu

Assessing and Verifying Task Utility in LLM-Powered Applications

May 03, 2024

Negar Arabzadeh, Siging Huo, Nikhil Mehta, Qinqyun Wu, Chi Wang, Ahmed Awadallah, Charles L. A. Clarke, Julia Kiseleva

Figure 1 for Assessing and Verifying Task Utility in LLM-Powered Applications

Figure 2 for Assessing and Verifying Task Utility in LLM-Powered Applications

Figure 3 for Assessing and Verifying Task Utility in LLM-Powered Applications

Figure 4 for Assessing and Verifying Task Utility in LLM-Powered Applications

Abstract:The rapid development of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents, assisting humans in their daily tasks. However, a significant gap remains in assessing to what extent LLM-powered applications genuinely enhance user experience and task execution efficiency. This highlights the need to verify utility of LLM-powered applications, particularly by ensuring alignment between the application's functionality and end-user needs. We introduce AgentEval, a novel framework designed to simplify the utility verification process by automatically proposing a set of criteria tailored to the unique purpose of any given application. This allows for a comprehensive assessment, quantifying the utility of an application against the suggested criteria. We present a comprehensive analysis of the effectiveness and robustness of AgentEval for two open source datasets including Math Problem solving and ALFWorld House-hold related tasks. For reproducibility purposes, we make the data, code and all the logs publicly available at https://bit.ly/3w3yKcS .

* arXiv admin note: text overlap with arXiv:2402.09015

Via

Access Paper or Ask Questions