Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Isaak Lim

Learning Fine-to-Coarse Cuboid Shape Abstraction

Feb 03, 2025

Gregor Kobsik, Morten Henkel, Yanjiang He, Victor Czech, Tim Elsner, Isaak Lim, Leif Kobbelt

Figure 1 for Learning Fine-to-Coarse Cuboid Shape Abstraction

Figure 2 for Learning Fine-to-Coarse Cuboid Shape Abstraction

Figure 3 for Learning Fine-to-Coarse Cuboid Shape Abstraction

Figure 4 for Learning Fine-to-Coarse Cuboid Shape Abstraction

Abstract:The abstraction of 3D objects with simple geometric primitives like cuboids allows to infer structural information from complex geometry. It is important for 3D shape understanding, structural analysis and geometric modeling. We introduce a novel fine-to-coarse unsupervised learning approach to abstract collections of 3D shapes. Our architectural design allows us to reduce the number of primitives from hundreds (fine reconstruction) to only a few (coarse abstraction) during training. This allows our network to optimize the reconstruction error and adhere to a user-specified number of primitives per shape while simultaneously learning a consistent structure across the whole collection of data. We achieve this through our abstraction loss formulation which increasingly penalizes redundant primitives. Furthermore, we introduce a reconstruction loss formulation to account not only for surface approximation but also volume preservation. Combining both contributions allows us to represent 3D shapes more precisely with fewer cuboid primitives than previous work. We evaluate our method on collections of man-made and humanoid shapes comparing with previous state-of-the-art learning methods on commonly used benchmarks. Our results confirm an improvement over previous cuboid-based shape abstraction techniques. Furthermore, we demonstrate our cuboid abstraction in downstream tasks like clustering, retrieval, and partial symmetry detection.

* 10 pages, 6 figures, 4 tables

Via

Access Paper or Ask Questions

Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation

Nov 15, 2024

Tim Elsner, Paula Usinger, Julius Nehring-Wirxel, Gregor Kobsik, Victor Czech, Yanjiang He, Isaak Lim, Leif Kobbelt

Figure 1 for Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation

Figure 2 for Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation

Figure 3 for Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation

Figure 4 for Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation

Abstract:In language processing, transformers benefit greatly from text being condensed. This is achieved through a larger vocabulary that captures word fragments instead of plain characters. This is often done with Byte Pair Encoding. In the context of images, tokenisation of visual data is usually limited to regular grids obtained from quantisation methods, without global content awareness. Our work improves tokenisation of visual data by bringing Byte Pair Encoding from 1D to multiple dimensions, as a complementary add-on to existing compression. We achieve this through counting constellations of token pairs and replacing the most frequent token pair with a newly introduced token. The multidimensionality only increases the computation time by a factor of 2 for images, making it applicable even to large datasets like ImageNet within minutes on consumer hardware. This is a lossless preprocessing step. Our evaluation shows improved training and inference performance of transformers on visual data achieved by compressing frequent constellations of tokens: The resulting sequences are shorter, with more uniformly distributed information content, e.g. condensing empty regions in an image into single tokens. As our experiments show, these condensed sequences are easier to process. We additionally introduce a strategy to amplify this compression further by clustering the vocabulary.

Via

Access Paper or Ask Questions

Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data

Jul 16, 2024

Tim Elsner, Paula Usinger, Victor Czech, Gregor Kobsik, Yanjiang He, Isaak Lim, Leif Kobbelt

Figure 1 for Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data

Figure 2 for Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data

Figure 3 for Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data

Figure 4 for Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data

Abstract:In quantised autoencoders, images are usually split into local patches, each encoded by one token. This representation is redundant in the sense that the same number of tokens is spend per region, regardless of the visual information content in that region. Adaptive discretisation schemes like quadtrees are applied to allocate tokens for patches with varying sizes, but this just varies the region of influence for a token which nevertheless remains a local descriptor. Modern architectures add an attention mechanism to the autoencoder which infuses some degree of global information into the local tokens. Despite the global context, tokens are still associated with a local image region. In contrast, our method is inspired by spectral decompositions which transform an input signal into a superposition of global frequencies. Taking the data-driven perspective, we learn custom basis functions corresponding to the codebook entries in our VQ-VAE setup. Furthermore, a decoder combines these basis functions in a non-linear fashion, going beyond the simple linear superposition of spectral decompositions. We can achieve this global description with an efficient transpose operation between features and channels and demonstrate our performance on compression.

Via

Access Paper or Ask Questions

Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Dec 13, 2023

Gregor Kobsik, Isaak Lim, Leif Kobbelt

Figure 1 for Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Figure 2 for Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Figure 3 for Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Figure 4 for Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Abstract:Symmetry detection, especially partial and extrinsic symmetry, is essential for various downstream tasks, like 3D geometry completion, segmentation, compression and structure-aware shape encoding or generation. In order to detect partial extrinsic symmetries, we propose to learn rotation, reflection, translation and scale invariant local shape features for geodesic point cloud patches via contrastive learning, which are robust across multiple classes and generalize over different datasets. We show that our approach is able to extract multiple valid solutions for this ambiguous problem. Furthermore, we introduce a novel benchmark test for partial extrinsic symmetry detection to evaluate our method. Lastly, we incorporate the detected symmetries together with a region growing algorithm to demonstrate a downstream task with the goal of computing symmetry-aware partitions of 3D shapes. To our knowledge, we are the first to propose a self-supervised data-driven method for partial extrinsic symmetry detection.

Via

Access Paper or Ask Questions

Adaptive Voronoi NeRFs

Mar 30, 2023

Tim Elsner, Victor Czech, Julia Berger, Zain Selman, Isaak Lim, Leif Kobbelt

Abstract:Neural Radiance Fields (NeRFs) learn to represent a 3D scene from just a set of registered images. Increasing sizes of a scene demands more complex functions, typically represented by neural networks, to capture all details. Training and inference then involves querying the neural network millions of times per image, which becomes impractically slow. Since such complex functions can be replaced by multiple simpler functions to improve speed, we show that a hierarchy of Voronoi diagrams is a suitable choice to partition the scene. By equipping each Voronoi cell with its own NeRF, our approach is able to quickly learn a scene representation. We propose an intuitive partitioning of the space that increases quality gains during training by distributing information evenly among the networks and avoids artifacts through a top-down adaptive refinement. Our framework is agnostic to the underlying NeRF method and easy to implement, which allows it to be applied to various NeRF variants for improved learning and rendering speeds.

Via

Access Paper or Ask Questions

Localized Latent Updates for Fine-Tuning Vision-Language Models

Dec 13, 2022

Moritz Ibing, Isaak Lim, Leif Kobbelt

Figure 1 for Localized Latent Updates for Fine-Tuning Vision-Language Models

Figure 2 for Localized Latent Updates for Fine-Tuning Vision-Language Models

Figure 3 for Localized Latent Updates for Fine-Tuning Vision-Language Models

Figure 4 for Localized Latent Updates for Fine-Tuning Vision-Language Models

Abstract:Although massive pre-trained vision-language models like CLIP show impressive generalization capabilities for many tasks, still it often remains necessary to fine-tune them for improved performance on specific datasets. When doing so, it is desirable that updating the model is fast and that the model does not lose its capabilities on data outside of the dataset, as is often the case with classical fine-tuning approaches. In this work we suggest a lightweight adapter, that only updates the models predictions close to seen datapoints. We demonstrate the effectiveness and speed of this relatively simple approach in the context of few-shot learning, where our results both on classes seen and unseen during training are comparable with or improve on the state of the art.

Via

Access Paper or Ask Questions

3D Shape Generation with Grid-based Implicit Functions

Jul 22, 2021

Moritz Ibing, Isaak Lim, Leif Kobbelt

Figure 1 for 3D Shape Generation with Grid-based Implicit Functions

Figure 2 for 3D Shape Generation with Grid-based Implicit Functions

Figure 3 for 3D Shape Generation with Grid-based Implicit Functions

Figure 4 for 3D Shape Generation with Grid-based Implicit Functions

Abstract:Previous approaches to generate shapes in a 3D setting train a GAN on the latent space of an autoencoder (AE). Even though this produces convincing results, it has two major shortcomings. As the GAN is limited to reproduce the dataset the AE was trained on, we cannot reuse a trained AE for novel data. Furthermore, it is difficult to add spatial supervision into the generation process, as the AE only gives us a global representation. To remedy these issues, we propose to train the GAN on grids (i.e. each cell covers a part of a shape). In this representation each cell is equipped with a latent vector provided by an AE. This localized representation enables more expressiveness (since the cell-based latent vectors can be combined in novel ways) as well as spatial control of the generation process (e.g. via bounding boxes). Our method outperforms the current state of the art on all established evaluation measures, proposed for quantitatively evaluating the generative capabilities of GANs. We show limitations of these measures and propose the adaptation of a robust criterion from statistical analysis as an alternative.

* CVPR 2021

Via

Access Paper or Ask Questions

A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Jun 27, 2019

Isaak Lim, Moritz Ibing, Leif Kobbelt

Figure 1 for A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Figure 2 for A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Figure 3 for A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Figure 4 for A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Abstract:Automatic synthesis of high quality 3D shapes is an ongoing and challenging area of research. While several data-driven methods have been proposed that make use of neural networks to generate 3D shapes, none of them reach the level of quality that deep learning synthesis approaches for images provide. In this work we present a method for a convolutional point cloud decoder/generator that makes use of recent advances in the domain of image synthesis. Namely, we use Adaptive Instance Normalization and offer an intuition on why it can improve training. Furthermore, we propose extensions to the minimization of the commonly used Chamfer distance for auto-encoding point clouds. In addition, we show that careful sampling is important both for the input geometry and in our point cloud generation process to improve results. The results are evaluated in an auto-encoding setup to offer both qualitative and quantitative analysis. The proposed decoder is validated by an extensive ablation study and is able to outperform current state of the art results in a number of experiments. We show the applicability of our method in the fields of point cloud upsampling, single view reconstruction, and shape synthesis.

* Computer Graphics Forum 38 (5), 2019
* Symposium on Geometry Processing 2019

Via

Access Paper or Ask Questions

A Simple Approach to Intrinsic Correspondence Learning on Unstructured 3D Meshes

Sep 26, 2018

Isaak Lim, Alexander Dielen, Marcel Campen, Leif Kobbelt

Figure 1 for A Simple Approach to Intrinsic Correspondence Learning on Unstructured 3D Meshes

Figure 2 for A Simple Approach to Intrinsic Correspondence Learning on Unstructured 3D Meshes

Figure 3 for A Simple Approach to Intrinsic Correspondence Learning on Unstructured 3D Meshes

Figure 4 for A Simple Approach to Intrinsic Correspondence Learning on Unstructured 3D Meshes

Abstract:The question of representation of 3D geometry is of vital importance when it comes to leveraging the recent advances in the field of machine learning for geometry processing tasks. For common unstructured surface meshes state-of-the-art methods rely on patch-based or mapping-based techniques that introduce resampling operations in order to encode neighborhood information in a structured and regular manner. We investigate whether such resampling can be avoided, and propose a simple and direct encoding approach. It does not only increase processing efficiency due to its simplicity - its direct nature also avoids any loss in data fidelity. To evaluate the proposed method, we perform a number of experiments in the challenging domain of intrinsic, non-rigid shape correspondence estimation. In comparisons to current methods we observe that our approach is able to achieve highly competitive results.

* Presented at the ECCV workshop on Geometry meets Deep Learning

Via

Access Paper or Ask Questions