Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:$ε$-VAE: Denoising as Visual Decoding

Oct 05, 2024

Long Zhao, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang, Boqing Gong, Hartwig Adam, Xuhui Jia, Ting Liu

Figure 1 for $ε$-VAE: Denoising as Visual Decoding

Figure 2 for $ε$-VAE: Denoising as Visual Decoding

Figure 3 for $ε$-VAE: Denoising as Visual Decoding

Figure 4 for $ε$-VAE: Denoising as Visual Decoding

Share this with someone who'll enjoy it:

Abstract:In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely on a traditional autoencoder framework, where the encoder compresses data into latent representations, and the decoder reconstructs the original input. In this work, we offer a new perspective by proposing denoising as decoding, shifting from single-step reconstruction to iterative refinement. Specifically, we replace the decoder with a diffusion process that iteratively refines noise to recover the original image, guided by the latents provided by the encoder. We evaluate our approach by assessing both reconstruction (rFID) and generation quality (FID), comparing it to state-of-the-art autoencoding approach. We hope this work offers new insights into integrating iterative generation and autoencoding for improved compression and generation.

View paper on

Share this with someone who'll enjoy it:

Title:$ε$-VAE: Denoising as Visual Decoding

Paper and Code