Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Haneol Jang

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Dec 13, 2024

Hyeonseok Lim, Dongjae Shin, Seohyun Song, Inho Won, Minjun Kim, Junghun Yuk, Haneol Jang, KyungTae Lim

Figure 1 for VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Figure 2 for VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Figure 3 for VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Figure 4 for VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Abstract:We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing evaluation datasets for external knowledge-based VQA, the proposed VLR-Bench includes five input passages. This allows testing of the ability to determine which passage is useful for answering a given query, a capability lacking in previous research. In this context, we constructed a dataset of 32,000 automatically generated instruction-following examples, which we denote as VLR-IF. This dataset is specifically designed to enhance the RAG capabilities of VLMs by enabling them to learn how to generate appropriate answers based on input passages. We evaluated the validity of the proposed benchmark and training data and verified its performance using the state-of-the-art Llama3-based VLM, the Llava-Llama-3 model. The proposed VLR-Bench and VLR-IF datasets are publicly available online.

* The 31st International Conference on Computational Linguistics (COLING 2025), 19 pages

Via

Access Paper or Ask Questions

Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

Jun 18, 2024

Hyojin Kim, Jiyoon Lee, Yonghyun Jeong, Haneol Jang, YoungJoon Yoo

Figure 1 for Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

Figure 2 for Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

Figure 3 for Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

Figure 4 for Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

Abstract:This paper presents a novel perspective for enhancing anti-spoofing performance in zero-shot data domain generalization. Unlike traditional image classification tasks, face anti-spoofing datasets display unique generalization characteristics, necessitating novel zero-shot data domain generalization. One step forward to the previous frame-wise spoofing prediction, we introduce a nuanced metric calculation that aggregates frame-level probabilities for a video-wise prediction, to tackle the gap between the reported frame-wise accuracy and instability in real-world use-case. This approach enables the quantification of bias and variance in model predictions, offering a more refined analysis of model generalization. Our investigation reveals that simply scaling up the backbone of models does not inherently improve the mentioned instability, leading us to propose an ensembled backbone method from a Bayesian perspective. The probabilistically ensembled backbone both improves model robustness measured from the proposed metric and spoofing accuracy, and also leverages the advantages of measuring uncertainty, allowing for enhanced sampling during training that contributes to model generalization across new datasets. We evaluate the proposed method from the benchmark OMIC dataset and also the public CelebA-Spoof and SiW-Mv2. Our final model outperforms existing state-of-the-art methods across the datasets, showcasing advancements in Bias, Variance, HTER, and AUC metrics.

* 10 pages with 4 figures, Accepted by CVPRW 2024

Via

Access Paper or Ask Questions

BOK-VQA: Bilingual Outside Knowledge-based Visual Question Answering via Graph Representation Pretraining

Jan 12, 2024

Minjun Kim, Seungwoo Song, Youhan Lee, Haneol Jang, Kyungtae Lim

Abstract:The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circumstances, the demand for multilingual evaluation of visual question answering (VQA) tasks, a representative task of multimodal systems, has increased. Accordingly, we propose a bilingual outside-knowledge VQA (BOK-VQA) dataset in this study that can be extended to multilingualism. The proposed data include 17K images, 17K question-answer pairs for both Korean and English and 280K instances of knowledge information related to question-answer content. We also present a framework that can effectively inject knowledge information into a VQA system by pretraining the knowledge information of BOK-VQA data in the form of graph embeddings. Finally, through in-depth analysis, we demonstrated the actual effect of the knowledge information contained in the constructed training data on VQA.

* Will be published at AAAI 2024

Via

Access Paper or Ask Questions