Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Hani Alomari

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Jun 26, 2025

Hani Alomari, Anushka Sivakumar, Andrew Zhang, Chris Thomas

Figure 1 for Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Figure 2 for Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Figure 3 for Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Figure 4 for Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Abstract:Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities. Set-based approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships. In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness. To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set. We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set. Our method achieves state-of-the-art performance on MS-COCO and Flickr30k without relying on external data.

* Accepted at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main)

Via

Access Paper or Ask Questions

ENTER: Event Based Interpretable Reasoning for VideoQA

Jan 24, 2025

Hammad Ayyubi, Junzhang Liu, Ali Asgarov, Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Zhecan Wang, Chia-Wei Tang, Hani Alomari, Md. Atabuzzaman, Xudong Lin(+3 more)

Figure 1 for ENTER: Event Based Interpretable Reasoning for VideoQA

Figure 2 for ENTER: Event Based Interpretable Reasoning for VideoQA

Figure 3 for ENTER: Event Based Interpretable Reasoning for VideoQA

Figure 4 for ENTER: Event Based Interpretable Reasoning for VideoQA

Abstract:In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-event relationships (temporal/causal/hierarchical) form the edges. This structured representation offers many benefits: 1) Interpretable VideoQA via generated code that parses event-graph; 2) Incorporation of contextual visual information in the reasoning process (code generation) via event graphs; 3) Robust VideoQA via Hierarchical Iterative Update of the event graphs. Existing interpretable VideoQA systems are often top-down, disregarding low-level visual information in the reasoning plan generation, and are brittle. While bottom-up approaches produce responses from visual data, they lack interpretability. Experimental results on NExT-QA, IntentQA, and EgoSchema demonstrate that not only does our method outperform existing top-down approaches while obtaining competitive performance against bottom-up approaches, but more importantly, offers superior interpretability and explainability in the reasoning process.

Via

Access Paper or Ask Questions