Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Xinyu Peng

Noise Conditional Variational Score Distillation

Jun 11, 2025

Xinyu Peng, Ziyang Zheng, Yaoming Wang, Han Li, Nuowen Kan, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong

Abstract:We propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function implicitly characterizes the score function of denoising posterior distributions. By integrating this insight into the Variational Score Distillation (VSD) framework, we enable scalable learning of generative denoisers capable of approximating samples from the denoising posterior distribution across a wide range of noise levels. The proposed generative denoisers exhibit desirable properties that allow fast generation while preserve the benefit of iterative refinement: (1) fast one-step generation through sampling from pure Gaussian noise at high noise levels; (2) improved sample quality by scaling the test-time compute with multi-step sampling; and (3) zero-shot probabilistic inference for flexible and controllable sampling. We evaluate NCVSD through extensive experiments, including class-conditional image generation and inverse problem solving. By scaling the test-time compute, our method outperforms teacher diffusion models and is on par with consistency models of larger sizes. Additionally, with significantly fewer NFEs than diffusion-based methods, we achieve record-breaking LPIPS on inverse problems.

Via

Access Paper or Ask Questions

Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Feb 03, 2024

Xinyu Peng, Ziyang Zheng, Wenrui Dai, Nuoqian Xiao, Chenglin Li, Junni Zou, Hongkai Xiong

Figure 1 for Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Figure 2 for Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Figure 3 for Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Figure 4 for Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Abstract:Recent diffusion models provide a promising zero-shot solution to noisy linear inverse problems without retraining for specific inverse problems. In this paper, we propose the first unified interpretation for existing zero-shot methods from the perspective of approximating the conditional posterior mean for the reverse diffusion process of conditional sampling. We reveal that recent methods are equivalent to making isotropic Gaussian approximations to intractable posterior distributions over clean images given diffused noisy images, with the only difference in the handcrafted design of isotropic posterior covariances. Inspired by this finding, we propose a general plug-and-play posterior covariance optimization based on maximum likelihood estimation to improve recent methods. To achieve optimal posterior covariance without retraining, we provide general solutions based on two approaches specifically designed to leverage pre-trained models with and without reverse covariances. Experimental results demonstrate that the proposed methods significantly enhance the overall performance or robustness to hyperparameters of recent methods. Code is available at https://github.com/xypeng9903/k-diffusion-inverse-problems

Via

Access Paper or Ask Questions

Drill the Cork of Information Bottleneck by Inputting the Most Important Data

May 15, 2021

Xinyu Peng, Jiawei Zhang, Fei-Yue Wang, Li Li

Figure 1 for Drill the Cork of Information Bottleneck by Inputting the Most Important Data

Figure 2 for Drill the Cork of Information Bottleneck by Inputting the Most Important Data

Figure 3 for Drill the Cork of Information Bottleneck by Inputting the Most Important Data

Figure 4 for Drill the Cork of Information Bottleneck by Inputting the Most Important Data

Abstract:Deep learning has become the most powerful machine learning tool in the last decade. However, how to efficiently train deep neural networks remains to be thoroughly solved. The widely used minibatch stochastic gradient descent (SGD) still needs to be accelerated. As a promising tool to better understand the learning dynamic of minibatch SGD, the information bottleneck (IB) theory claims that the optimization process consists of an initial fitting phase and the following compression phase. Based on this principle, we further study typicality sampling, an efficient data selection method, and propose a new explanation of how it helps accelerate the training process of the deep networks. We show that the fitting phase depicted in the IB theory will be boosted with a high signal-to-noise ratio of gradient approximation if the typicality sampling is appropriately adopted. Furthermore, this finding also implies that the prior information of the training set is critical to the optimization process and the better use of the most important data can help the information flow through the bottleneck faster. Both theoretical analysis and experimental results on synthetic and real-world datasets demonstrate our conclusions.

* 11 pages, to be published in IEEE Transactions on Neural Networks and Learning Systems

Via

Access Paper or Ask Questions

Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling

Mar 11, 2019

Xinyu Peng, Li Li, Fei-Yue Wang

Figure 1 for Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling

Figure 2 for Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling

Figure 3 for Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling

Figure 4 for Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling

Abstract:Machine learning, especially deep neural networks, has been rapidly developed in fields including computer vision, speech recognition and reinforcement learning. Although Mini-batch SGD is one of the most popular stochastic optimization methods in training deep networks, it shows a slow convergence rate due to the large noise in gradient approximation. In this paper, we attempt to remedy this problem by building more efficient batch selection method based on typicality sampling, which reduces the error of gradient estimation in conventional Minibatch SGD. We analyze the convergence rate of the resulting typical batch SGD algorithm and compare convergence properties between Minibatch SGD and the algorithm. Experimental results demonstrate that our batch selection scheme works well and more complex Minibatch SGD variants can benefit from the proposed batch selection strategy.

* 10 pages, 4 figures, for journal

Via

Access Paper or Ask Questions