Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Juliette Marrie

CASA: Cross-Attention via Self-Attention for Efficient Vision-Language Fusion

Dec 22, 2025

Moritz Böhle, Amélie Royer, Juliette Marrie, Edouard Grave, Patrick Pérez

Abstract:Vision-language models (VLMs) are commonly trained by inserting image tokens from a pretrained vision encoder into the textual stream of a language model. This allows text and image information to fully attend to one another within the model, but becomes extremely costly for high-resolution images, long conversations, or streaming videos, both in memory and compute. VLMs leveraging cross-attention are an efficient alternative to token insertion but exhibit a clear performance gap, in particular on tasks involving fine-grained visual details. We find that a key to improving such models is to also enable local text-to-text interaction in the dedicated cross-attention layers. Building on this, we propose CASA, Cross-Attention via Self-Attention, a simple and efficient paradigm which substantially reduces the gap with full token insertion on common image understanding benchmarks, while enjoying the same scalability as cross-attention models when applied to long-context multimodal tasks such as streaming video captioning. For samples and code, please see our project page at https://kyutai.org/casa .

Via

Access Paper or Ask Questions

PanSt3R: Multi-view Consistent Panoptic Segmentation

Jun 26, 2025

Lojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld, Boris Chidlovskii, Jerome Revaud, Gabriela Csurka

Abstract:Panoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing approaches typically leverage off-the-shelf models to extract per-frame 2D panoptic segmentations, before optimizing an implicit geometric representation (often based on NeRF) to integrate and fuse the 2D predictions. We argue that relying on 2D panoptic segmentation for a problem inherently 3D and multi-view is likely suboptimal as it fails to leverage the full potential of spatial relationships across views. In addition to requiring camera parameters, these approaches also necessitate computationally expensive test-time optimization for each scene. Instead, in this work, we propose a unified and integrated approach PanSt3R, which eliminates the need for test-time optimization by jointly predicting 3D geometry and multi-view panoptic segmentation in a single forward pass. Our approach builds upon recent advances in 3D reconstruction, specifically upon MUSt3R, a scalable multi-view version of DUSt3R, and enhances it with semantic awareness and multi-view panoptic segmentation capabilities. We additionally revisit the standard post-processing mask merging procedure and introduce a more principled approach for multi-view segmentation. We also introduce a simple method for generating novel-view predictions based on the predictions of PanSt3R and vanilla 3DGS. Overall, the proposed PanSt3R is conceptually simple, yet fast and scalable, and achieves state-of-the-art performance on several benchmarks, while being orders of magnitude faster than existing methods.

* Accepted at ICCV 2025

Via

Access Paper or Ask Questions

LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

Oct 18, 2024

Juliette Marrie, Romain Ménégaux, Michael Arbel, Diane Larlus, Julien Mairal

Figure 1 for LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

Figure 2 for LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

Figure 3 for LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

Figure 4 for LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

Abstract:We address the task of uplifting visual features or semantic masks from 2D vision models to 3D scenes represented by Gaussian Splatting. Whereas common approaches rely on iterative optimization-based procedures, we show that a simple yet effective aggregation technique yields excellent results. Applied to semantic masks from Segment Anything (SAM), our uplifting approach leads to segmentation quality comparable to the state of the art. We then extend this method to generic DINOv2 features, integrating 3D scene geometry through graph diffusion, and achieve competitive segmentation results despite DINOv2 not being trained on millions of annotated masks like SAM.

Via

Access Paper or Ask Questions

On Good Practices for Task-Specific Distillation of Large Pretrained Models

Feb 17, 2024

Juliette Marrie, Michael Arbel, Julien Mairal, Diane Larlus

Figure 1 for On Good Practices for Task-Specific Distillation of Large Pretrained Models

Figure 2 for On Good Practices for Task-Specific Distillation of Large Pretrained Models

Figure 3 for On Good Practices for Task-Specific Distillation of Large Pretrained Models

Figure 4 for On Good Practices for Task-Specific Distillation of Large Pretrained Models

Abstract:Large pretrained visual models exhibit remarkable generalization across diverse recognition tasks. Yet, real-world applications often demand compact models tailored to specific problems. Variants of knowledge distillation have been devised for such a purpose, enabling task-specific compact models (the students) to learn from a generic large pretrained one (the teacher). In this paper, we show that the excellent robustness and versatility of recent pretrained models challenge common practices established in the literature, calling for a new set of optimal guidelines for task-specific distillation. To address the lack of samples in downstream tasks, we also show that a variant of Mixup based on stable diffusion complements standard data augmentation. This strategy eliminates the need for engineered text prompts and improves distillation of generic models into streamlined specialized networks.

Via

Access Paper or Ask Questions

SLACK: Stable Learning of Augmentations with Cold-start and KL regularization

Jun 16, 2023

Juliette Marrie, Michael Arbel, Diane Larlus, Julien Mairal

Figure 1 for SLACK: Stable Learning of Augmentations with Cold-start and KL regularization

Figure 2 for SLACK: Stable Learning of Augmentations with Cold-start and KL regularization

Figure 3 for SLACK: Stable Learning of Augmentations with Cold-start and KL regularization

Figure 4 for SLACK: Stable Learning of Augmentations with Cold-start and KL regularization

Abstract:Data augmentation is known to improve the generalization capabilities of neural networks, provided that the set of transformations is chosen with care, a selection often performed manually. Automatic data augmentation aims at automating this process. However, most recent approaches still rely on some prior information; they start from a small pool of manually-selected default transformations that are either used to pretrain the network or forced to be part of the policy learned by the automatic data augmentation algorithm. In this paper, we propose to directly learn the augmentation policy without leveraging such prior knowledge. The resulting bilevel optimization problem becomes more challenging due to the larger search space and the inherent instability of bilevel optimization algorithms. To mitigate these issues (i) we follow a successive cold-start strategy with a Kullback-Leibler regularization, and (ii) we parameterize magnitudes as continuous distributions. Our approach leads to competitive results on standard benchmarks despite a more challenging setting, and generalizes beyond natural images.

* Accepted to CVPR 2023

Via

Access Paper or Ask Questions