Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Apr 08, 2025

Peerat Limkonchotiwat, Kanruethai Masuk, Surapon Nonesung, Chalermpun Mai-On, Sarana Nutanong, Wuttikorn Ponwitayarat, Potsawee Manakul

Figure 1 for Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Figure 2 for Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Figure 3 for Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Figure 4 for Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Share this with someone who'll enjoy it:

Abstract:Large language models show promising results in various NLP tasks. Despite these successes, the robustness and consistency of LLMs in underrepresented languages remain largely unexplored, especially concerning local dialects. Existing benchmarks also focus on main dialects, neglecting LLMs' ability on local dialect texts. In this paper, we introduce a Thai local dialect benchmark covering Northern (Lanna), Northeastern (Isan), and Southern (Dambro) Thai, evaluating LLMs on five NLP tasks: summarization, question answering, translation, conversation, and food-related tasks. Furthermore, we propose a human evaluation guideline and metric for Thai local dialects to assess generation fluency and dialect-specific accuracy. Results show that LLM performance declines significantly in local Thai dialects compared to standard Thai, with only proprietary models like GPT-4o and Gemini2 demonstrating some fluency

* Datasets and codes are available at https://github.com/mrpeerat/Thai_local_benchmark

View paper on

Share this with someone who'll enjoy it:

Title:Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation

Paper and Code