Picture for Davide Caffagni

Davide Caffagni

Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization

Add code
Aug 26, 2024
Viaarxiv icon

Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Add code
Apr 23, 2024
Figure 1 for Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Figure 2 for Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Figure 3 for Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Figure 4 for Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Viaarxiv icon

The (R)Evolution of Multimodal Large Language Models: A Survey

Add code
Feb 19, 2024
Viaarxiv icon