Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Runzhong Zhang

Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation

Feb 06, 2025

Yang Chen, Yueqi Duan, Runzhong Zhang, Yap-Peng Tan

Figure 1 for Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation

Figure 2 for Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation

Figure 3 for Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation

Figure 4 for Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation

Abstract:In this paper, we propose an adaptive margin contrastive learning method for 3D point cloud semantic segmentation, namely AMContrast3D. Most existing methods use equally penalized objectives, which ignore per-point ambiguities and less discriminated features stemming from transition regions. However, as highly ambiguous points may be indistinguishable even for humans, their manually annotated labels are less reliable, and hard constraints over these points would lead to sub-optimal models. To address this, we design adaptive objectives for individual points based on their ambiguity levels, aiming to ensure the correctness of low-ambiguity points while allowing mistakes for high-ambiguity points. Specifically, we first estimate ambiguities based on position embeddings. Then, we develop a margin generator to shift decision boundaries for contrastive feature embeddings, so margins are narrowed due to increasing ambiguities with even negative margins for extremely high-ambiguity points. Experimental results on large-scale datasets, S3DIS and ScanNet, demonstrate that our method outperforms state-of-the-art methods.

* 2024 IEEE International Conference on Multimedia and Expo (ICME)

Via

Access Paper or Ask Questions

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Oct 01, 2024

Chen Cai, Zheng Wang, Jianjun Gao, Wenyang Liu, Ye Lu, Runzhong Zhang, Kim-Hui Yap

Figure 1 for Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Figure 2 for Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Figure 3 for Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Figure 4 for Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Abstract:In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting. To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting. These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research. Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14\% accuracy on NExT-QA and 71.24\% accuracy on DramaQA, highlighting its practical relevance and effectiveness.

* Accepted by main EMNLP 2024

Via

Access Paper or Ask Questions

Video sentence grounding with temporally global textual knowledge

Apr 21, 2024

Cai Chen, Runzhong Zhang, Jianjun Gao, Kejun Wu, Kim-Hui Yap, Yi Wang

Abstract:Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multi-modal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method.

Via

Access Paper or Ask Questions