Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Xiongtao Zhou

MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Oct 18, 2024

Xiongtao Zhou, Jie He, Lanyu Chen, jingyu li, Haojing Chen, Victor Gutierrez Basulto, Jeff Z. Pan, Hanjie Chen

Figure 1 for MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Figure 2 for MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Figure 3 for MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Figure 4 for MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Abstract:Multimodal Chain of Thought (MCoT) is a popular prompting strategy for improving the performance of multimodal large language models (MLLMs) across a range of complex reasoning tasks. Despite its popularity, there is a notable absence of automated methods for evaluating the quality of reasoning steps in MCoT. To address this gap, we propose Multimodal Chain-of-Thought Evaluation (MiCEval), a framework designed to assess the correctness of reasoning chains by evaluating the quality of both the description and each reasoning step. The evaluation of the description component focuses on the accuracy of the image descriptions, while the reasoning step evaluates the quality of each step as it is conditionally generated based on the preceding steps. MiCEval is built upon a fine-grained dataset with annotations that rate each step according to correctness, relevance, and informativeness. Extensive experiments on four state-of-the-art MLLMs show that step-wise evaluations using MiCEval align more closely with human judgments compared to existing methods based on cosine similarity or fine-tuning approaches. MiCEval datasets and code can be found in https://github.com/alenai97/MiCEval.

* 40 pages

Via

Access Paper or Ask Questions

An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Jun 07, 2024

Xiongtao Zhou, Jie He, Yuhua Ke, Guangyao Zhu, Víctor Gutiérrez-Basulto, Jeff Z. Pan

Figure 1 for An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Figure 2 for An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Figure 3 for An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Figure 4 for An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Abstract:Multimodal large language models (MLLMs) fine-tuned with multimodal instruction datasets have demonstrated remarkable capabilities in multimodal tasks. However, fine-tuning all parameters of MLLMs has become challenging as they usually contain billions of parameters. To address this issue, we study parameter-efficient fine-tuning (PEFT) methods for MLLMs. We aim to identify effective methods for enhancing the performance of MLLMs in scenarios where only a limited number of parameters are trained. This paper conducts empirical studies using four popular PEFT methods to fine-tune the LLM component of open-source MLLMs. We present a comprehensive analysis that encompasses various aspects, including the impact of PEFT methods on various models, parameters and location of the PEFT module, size of fine-tuning data, model stability based on PEFT methods, MLLM's generalization, and hallucination. We evaluated four PEFT methods on seven datasets from two different categories: unseen and seen datasets. Across all experiments, we show that the adapter is the best-performing PEFT method. At the same time, fine-tuning the connector layers leads to improved performance in most MLLMs. Code and data are available at https://github.com/alenai97/PEFT-MLLM.git.

* ACL finding 2024

Via

Access Paper or Ask Questions