Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Felix Jimenez

Probabilistic Skip Connections for Deterministic Uncertainty Quantification in Deep Neural Networks

Jan 08, 2025

Felix Jimenez, Matthias Katzfuss

Figure 1 for Probabilistic Skip Connections for Deterministic Uncertainty Quantification in Deep Neural Networks

Figure 2 for Probabilistic Skip Connections for Deterministic Uncertainty Quantification in Deep Neural Networks

Figure 3 for Probabilistic Skip Connections for Deterministic Uncertainty Quantification in Deep Neural Networks

Figure 4 for Probabilistic Skip Connections for Deterministic Uncertainty Quantification in Deep Neural Networks

Abstract:Deterministic uncertainty quantification (UQ) in deep learning aims to estimate uncertainty with a single pass through a network by leveraging outputs from the network's feature extractor. Existing methods require that the feature extractor be both sensitive and smooth, ensuring meaningful input changes produce meaningful changes in feature vectors. Smoothness enables generalization, while sensitivity prevents feature collapse, where distinct inputs are mapped to identical feature vectors. To meet these requirements, current deterministic methods often retrain networks with spectral normalization. Instead of modifying training, we propose using measures of neural collapse to identify an existing intermediate layer that is both sensitive and smooth. We then fit a probabilistic model to the feature vector of this intermediate layer, which we call a probabilistic skip connection (PSC). Through empirical analysis, we explore the impact of spectral normalization on neural collapse and demonstrate that PSCs can effectively disentangle aleatoric and epistemic uncertainty. Additionally, we show that PSCs achieve uncertainty quantification and out-of-distribution (OOD) detection performance that matches or exceeds existing single-pass methods requiring training modifications. By retrofitting existing models, PSCs enable high-quality UQ and OOD capabilities without retraining.

* 15 pages, 9 figures

Via

Access Paper or Ask Questions

Vecchia Gaussian Process Ensembles on Internal Representations of Deep Neural Networks

May 26, 2023

Felix Jimenez, Matthias Katzfuss

Abstract:For regression tasks, standard Gaussian processes (GPs) provide natural uncertainty quantification, while deep neural networks (DNNs) excel at representation learning. We propose to synergistically combine these two approaches in a hybrid method consisting of an ensemble of GPs built on the output of hidden layers of a DNN. GP scalability is achieved via Vecchia approximations that exploit nearest-neighbor conditional independence. The resulting deep Vecchia ensemble not only imbues the DNN with uncertainty quantification but can also provide more accurate and robust predictions. We demonstrate the utility of our model on several datasets and carry out experiments to understand the inner workings of the proposed method.

* 16 pages, 7 figures

Via

Access Paper or Ask Questions

Variational sparse inverse Cholesky approximation for latent Gaussian processes via double Kullback-Leibler minimization

Jan 30, 2023

Jian Cao, Myeongjong Kang, Felix Jimenez, Huiyan Sang, Florian Schafer, Matthias Katzfuss

Figure 1 for Variational sparse inverse Cholesky approximation for latent Gaussian processes via double Kullback-Leibler minimization

Figure 2 for Variational sparse inverse Cholesky approximation for latent Gaussian processes via double Kullback-Leibler minimization

Figure 3 for Variational sparse inverse Cholesky approximation for latent Gaussian processes via double Kullback-Leibler minimization

Figure 4 for Variational sparse inverse Cholesky approximation for latent Gaussian processes via double Kullback-Leibler minimization

Abstract:To achieve scalable and accurate inference for latent Gaussian processes, we propose a variational approximation based on a family of Gaussian distributions whose covariance matrices have sparse inverse Cholesky (SIC) factors. We combine this variational approximation of the posterior with a similar and efficient SIC-restricted Kullback-Leibler-optimal approximation of the prior. We then focus on a particular SIC ordering and nearest-neighbor-based sparsity pattern resulting in highly accurate prior and posterior approximations. For this setting, our variational approximation can be computed via stochastic gradient descent in polylogarithmic time per iteration. We provide numerical comparisons showing that the proposed double-Kullback-Leibler-optimal Gaussian-process approximation (DKLGP) can sometimes be vastly more accurate than alternative approaches such as inducing-point and mean-field approximations at similar computational complexity.

Via

Access Paper or Ask Questions

Scalable Bayesian Optimization Using Vecchia Approximations of Gaussian Processes

Mar 02, 2022

Felix Jimenez, Matthias Katzfuss

Figure 1 for Scalable Bayesian Optimization Using Vecchia Approximations of Gaussian Processes

Figure 2 for Scalable Bayesian Optimization Using Vecchia Approximations of Gaussian Processes

Figure 3 for Scalable Bayesian Optimization Using Vecchia Approximations of Gaussian Processes

Figure 4 for Scalable Bayesian Optimization Using Vecchia Approximations of Gaussian Processes

Abstract:Bayesian optimization is a technique for optimizing black-box target functions. At the core of Bayesian optimization is a surrogate model that predicts the output of the target function at previously unseen inputs to facilitate the selection of promising input values. Gaussian processes (GPs) are commonly used as surrogate models but are known to scale poorly with the number of observations. We adapt the Vecchia approximation, a popular GP approximation from spatial statistics, to enable scalable high-dimensional Bayesian optimization. We develop several improvements and extensions, including training warped GPs using mini-batch gradient descent, approximate neighbor search, and selecting multiple input values in parallel. We focus on the use of our warped Vecchia GP in trust-region Bayesian optimization via Thompson sampling. On several test functions and on two reinforcement-learning problems, our methods compared favorably to the state of the art.

Via

Access Paper or Ask Questions