Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Zhentao He

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Jan 13, 2025

Ziheng Wu, Zhenghao Chen, Ruipu Luo, Can Zhang, Yuan Gao, Zhentao He, Xian Wang, Haoran Lin, Minghui Qiu

Figure 1 for Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Figure 2 for Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Figure 3 for Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Figure 4 for Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Abstract:Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introduce Valley2, a novel multimodal large language model designed to enhance performance across all domains and extend the boundaries of practical applications in e-commerce and short video scenarios. Notably, Valley2 achieves state-of-the-art (SOTA) performance on e-commerce benchmarks, surpassing open-source models of similar size by a large margin (79.66 vs. 72.76). Additionally, Valley2 ranks second on the OpenCompass leaderboard among models with fewer than 10B parameters, with an impressive average score of 67.4. The code and model weights are open-sourced at https://github.com/bytedance/Valley.

Via

Access Paper or Ask Questions

PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network

Aug 02, 2024

Changqun Xia, Chenxi Xie, Zhentao He, Tianshu Yu, Jia Li

Figure 1 for PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network

Figure 2 for PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network

Figure 3 for PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network

Figure 4 for PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network

Abstract:We present an advanced study on more challenging high-resolution salient object detection (HRSOD) from both dataset and network framework perspectives. To compensate for the lack of HRSOD dataset, we thoughtfully collect a large-scale high resolution salient object detection dataset, called UHRSD, containing 5,920 images from real-world complex scenarios at 4K-8K resolutions. All the images are finely annotated in pixel-level, far exceeding previous low-resolution SOD datasets. Aiming at overcoming the contradiction between the sampling depth and the receptive field size in the past methods, we propose a novel one-stage framework for HR-SOD task using pyramid grafting mechanism. In general, transformer-based and CNN-based backbones are adopted to extract features from different resolution images independently and then these features are grafted from transformer branch to CNN branch. An attention-based Cross-Model Grafting Module (CMGM) is proposed to enable CNN branch to combine broken detailed information more holistically, guided by different source feature during decoding process. Moreover, we design an Attention Guided Loss (AGL) to explicitly supervise the attention matrix generated by CMGM to help the network better interact with the attention from different branches. Comprehensive experiments on UHRSD and widely-used SOD datasets demonstrate that our method can simultaneously locate salient object and preserve rich details, outperforming state-of-the-art methods. To verify the generalization ability of the proposed framework, we apply it to the camouflaged object detection (COD) task. Notably, our method performs superior to most state-of-the-art COD methods without bells and whistles.

Via

Access Paper or Ask Questions