Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Satya Narayan Shukla

Transfer between Modalities with MetaQueries

Apr 08, 2025

Xichen Pan, Satya Narayan Shukla, Aashu Singh, Zhuokai Zhao, Shlok Kumar Mishra, Jialiang Wang, Zhiyang Xu, Jiuhai Chen, Kunpeng Li, Felix Juefei-Xu(+2 more)

Abstract:Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often demands complex training recipes and careful data balancing. We introduce MetaQueries, a set of learnable queries that act as an efficient interface between autoregressive multimodal LLMs (MLLMs) and diffusion models. MetaQueries connects the MLLM's latents to the diffusion decoder, enabling knowledge-augmented image generation by leveraging the MLLM's deep understanding and reasoning capabilities. Our method simplifies training, requiring only paired image-caption data and standard diffusion objectives. Notably, this transfer is effective even when the MLLM backbone remains frozen, thereby preserving its state-of-the-art multimodal understanding capabilities while achieving strong generative performance. Additionally, our method is flexible and can be easily instruction-tuned for advanced applications such as image editing and subject-driven generation.

* Project Page: https://xichenpan.com/metaquery

Via

Access Paper or Ask Questions

CompCap: Improving Multimodal Large Language Models with Composite Captions

Dec 06, 2024

Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh, Qifan Wang, David Yang, ShengYun Peng, Hanchao Yu, Shen Yan, Xuewen Zhang(+1 more)

Figure 1 for CompCap: Improving Multimodal Large Language Models with Composite Captions

Figure 2 for CompCap: Improving Multimodal Large Language Models with Composite Captions

Figure 3 for CompCap: Improving Multimodal Large Language Models with Composite Captions

Figure 4 for CompCap: Improving Multimodal Large Language Models with Composite Captions

Abstract:How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively.

Via

Access Paper or Ask Questions

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Apr 11, 2024

Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, Tsung-Yu Lin

Figure 1 for Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Figure 2 for Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Figure 3 for Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Figure 4 for Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Abstract:Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing V-LLMs (e.g. BLIP-2, LLaVA) demonstrate weak spatial reasoning and localization awareness. Despite generating highly descriptive and elaborate textual answers, these models fail at simple tasks like distinguishing a left vs right location. In this work, we explore how image-space coordinate based instruction fine-tuning objectives could inject spatial awareness into V-LLMs. We discover optimal coordinate representations, data-efficient instruction fine-tuning objectives, and pseudo-data generation strategies that lead to improved spatial awareness in V-LLMs. Additionally, our resulting model improves VQA across image and video domains, reduces undesired hallucination, and generates better contextual object descriptions. Experiments across 5 vision-language tasks involving 14 different datasets establish the clear performance improvements achieved by our proposed framework.

Via

Access Paper or Ask Questions

Universal Pyramid Adversarial Training for Improved ViT Performance

Dec 26, 2023

Ping-yeh Chiang, Yipin Zhou, Omid Poursaeed, Satya Narayan Shukla, Ashish Shah, Tom Goldstein, Ser-Nam Lim

Figure 1 for Universal Pyramid Adversarial Training for Improved ViT Performance

Figure 2 for Universal Pyramid Adversarial Training for Improved ViT Performance

Figure 3 for Universal Pyramid Adversarial Training for Improved ViT Performance

Figure 4 for Universal Pyramid Adversarial Training for Improved ViT Performance

Abstract:Recently, Pyramid Adversarial training (Herrmann et al., 2022) has been shown to be very effective for improving clean accuracy and distribution-shift robustness of vision transformers. However, due to the iterative nature of adversarial training, the technique is up to 7 times more expensive than standard training. To make the method more efficient, we propose Universal Pyramid Adversarial training, where we learn a single pyramid adversarial pattern shared across the whole dataset instead of the sample-wise patterns. With our proposed technique, we decrease the computational cost of Pyramid Adversarial training by up to 70% while retaining the majority of its benefit on clean performance and distribution-shift robustness. In addition, to the best of our knowledge, we are also the first to find that universal adversarial training can be leveraged to improve clean model performance.

Via

Access Paper or Ask Questions

Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding

Sep 20, 2023

Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed, Pengchuan Zhang, Ashish Shah, Sernam Lim

Abstract:While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length. A common approach to process long videos is applying a short-form video model over uniformly sampled clips of fixed temporal length and aggregating the outputs. This approach neglects the underlying nature of long videos since fixed-length clips are often redundant or uninformative. In this paper, we aim to provide a generic and adaptive sampling approach for long-form videos in lieu of the de facto uniform sampling. Viewing videos as semantically consistent segments, we formulate a task-agnostic, unsupervised, and scalable approach based on Kernel Temporal Segmentation (KTS) for sampling and tokenizing long videos. We evaluate our method on long-form video understanding tasks such as video classification and temporal action localization, showing consistent gains over existing approaches and achieving state-of-the-art performance on long-form video modeling.

Via

Access Paper or Ask Questions

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Aug 31, 2023

Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa

Figure 1 for The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Figure 2 for The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Figure 3 for The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Figure 4 for The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Abstract:We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this dataset enables the evaluation of text models in high-, medium-, and low-resource languages. Each question is based on a short passage from the Flores-200 dataset and has four multiple-choice answers. The questions were carefully curated to discriminate between models with different levels of general language comprehension. The English dataset on its own proves difficult enough to challenge state-of-the-art language models. Being fully parallel, this dataset enables direct comparison of model performance across all languages. We use this dataset to evaluate the capabilities of multilingual masked language models (MLMs) and large language models (LLMs). We present extensive results and find that despite significant cross-lingual transfer in English-centric LLMs, much smaller MLMs pretrained on balanced multilingual data still understand far more languages. We also observe that larger vocabulary size and conscious vocabulary construction correlate with better performance on low-resource languages. Overall, Belebele opens up new avenues for evaluating and analyzing the multilingual capabilities of NLP systems.

* 27 pages, 13 figures

Via

Access Paper or Ask Questions

Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series

Jul 23, 2021

Satya Narayan Shukla, Benjamin M. Marlin

Figure 1 for Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series

Figure 2 for Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series

Figure 3 for Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series

Figure 4 for Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series

Abstract:Irregularly sampled time series commonly occur in several domains where they present a significant challenge to standard deep learning models. In this paper, we propose a new deep learning framework for probabilistic interpolation of irregularly sampled time series that we call the Heteroscedastic Temporal Variational Autoencoder (HeTVAE). HeTVAE includes a novel input layer to encode information about input observation sparsity, a temporal VAE architecture to propagate uncertainty due to input sparsity, and a heteroscedastic output layer to enable variable uncertainty in output interpolations. Our results show that the proposed architecture is better able to reflect variable uncertainty through time due to sparse and irregular sampling than a range of baseline and traditional models, as well as recently proposed deep latent variable models that use homoscedastic output layers.

Via

Access Paper or Ask Questions

Multi-Time Attention Networks for Irregularly Sampled Time Series

Jan 25, 2021

Satya Narayan Shukla, Benjamin M. Marlin

Figure 1 for Multi-Time Attention Networks for Irregularly Sampled Time Series

Figure 2 for Multi-Time Attention Networks for Irregularly Sampled Time Series

Abstract:Irregular sampling occurs in many time series modeling applications where it presents a significant challenge to standard deep learning models. This work is motivated by the analysis of physiological time series data in electronic health records, which are sparse, irregularly sampled, and multivariate. In this paper, we propose a new deep learning framework for this setting that we call Multi-Time Attention Networks. Multi-Time Attention Networks learn an embedding of continuous-time values and use an attention mechanism to produce a fixed-length representation of a time series containing a variable number of observations. We investigate the performance of our framework on interpolation and classification tasks using multiple datasets. Our results show that our approach performs as well or better than a range of baseline and recently proposed models while offering significantly faster training times than current state-of-the-art methods.

* Accepted at International Conference on Learning Representations (ICLR) 2021

Via

Access Paper or Ask Questions

A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series

Jan 05, 2021

Satya Narayan Shukla, Benjamin M. Marlin

Figure 1 for A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series

Figure 2 for A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series

Figure 3 for A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series

Figure 4 for A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series

Abstract:Irregularly sampled time series data arise naturally in many application domains including biology, ecology, climate science, astronomy, and health. Such data represent fundamental challenges to many classical models from machine learning and statistics due to the presence of non-uniform intervals between observations. However, there has been significant progress within the machine learning community over the last decade on developing specialized models and architectures for learning from irregularly sampled univariate and multivariate time series data. In this survey, we first describe several axes along which approaches to learning from irregularly sampled time series differ including what data representations they are based on, what modeling primitives they leverage to deal with the fundamental problem of irregular sampling, and what inference tasks they are designed to perform. We then survey the recent literature organized primarily along the axis of modeling primitives. We describe approaches based on temporal discretization, interpolation, recurrence, attention and structural invariance. We discuss similarities and differences between approaches and highlight primary strengths and weaknesses.

* Presented at NeurIPS 2020 Workshop: ML Retrospectives, Surveys & Meta-Analyses (ML-RSA)

Via

Access Paper or Ask Questions

Gaussian MRF Covariance Modeling for Efficient Black-Box Adversarial Attacks

Oct 08, 2020

Anit Kumar Sahu, Satya Narayan Shukla, J. Zico Kolter

Figure 1 for Gaussian MRF Covariance Modeling for Efficient Black-Box Adversarial Attacks

Figure 2 for Gaussian MRF Covariance Modeling for Efficient Black-Box Adversarial Attacks

Figure 3 for Gaussian MRF Covariance Modeling for Efficient Black-Box Adversarial Attacks

Figure 4 for Gaussian MRF Covariance Modeling for Efficient Black-Box Adversarial Attacks

Abstract:We study the problem of generating adversarial examples in a black-box setting, where we only have access to a zeroth order oracle, providing us with loss function evaluations. Although this setting has been investigated in previous work, most past approaches using zeroth order optimization implicitly assume that the gradients of the loss function with respect to the input images are \emph{unstructured}. In this work, we show that in fact substantial correlations exist within these gradients, and we propose to capture these correlations via a Gaussian Markov random field (GMRF). Given the intractability of the explicit covariance structure of the MRF, we show that the covariance structure can be efficiently represented using the Fast Fourier Transform (FFT), along with low-rank updates to perform exact posterior estimation under this model. We use this modeling technique to find fast one-step adversarial attacks, akin to a black-box version of the Fast Gradient Sign Method~(FGSM), and show that the method uses fewer queries and achieves higher attack success rates than the current state of the art. We also highlight the general applicability of this gradient modeling setup.

Via

Access Paper or Ask Questions