Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Tobias Weis

Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures

Dec 14, 2018

Martin Mundt, Sagnik Majumder, Tobias Weis, Visvanathan Ramesh

Figure 1 for Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures

Figure 2 for Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures

Figure 3 for Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures

Abstract:We characterize convolutional neural networks with respect to the relative amount of features per layer. Using a skew normal distribution as a parametrized framework, we investigate the common assumption of monotonously increasing feature-counts with higher layers of architecture designs. Our evaluation on models with VGG-type layers on the MNIST, Fashion-MNIST and CIFAR-10 image classification benchmarks provides evidence that motivates rethinking of our common assumption: architectures that favor larger early layers seem to yield better accuracy.

* Accepted at the Critiquing and Correcting Trends in Machine Learning (CRACT) Workshop at the 32nd Conference on Neural Information Processing Systems (NeurIPS 2018)

Via

Access Paper or Ask Questions

Building effective deep neural network architectures one feature at a time

Oct 19, 2017

Martin Mundt, Tobias Weis, Kishore Konda, Visvanathan Ramesh

Figure 1 for Building effective deep neural network architectures one feature at a time

Figure 2 for Building effective deep neural network architectures one feature at a time

Figure 3 for Building effective deep neural network architectures one feature at a time

Figure 4 for Building effective deep neural network architectures one feature at a time

Abstract:Successful training of convolutional neural networks is often associated with sufficiently deep architectures composed of high amounts of features. These networks typically rely on a variety of regularization and pruning techniques to converge to less redundant states. We introduce a novel bottom-up approach to expand representations in fixed-depth architectures. These architectures start from just a single feature per layer and greedily increase width of individual layers to attain effective representational capacities needed for a specific task. While network growth can rely on a family of metrics, we propose a computationally efficient version based on feature time evolution and demonstrate its potency in determining feature importance and a networks' effective capacity. We demonstrate how automatically expanded architectures converge to similar topologies that benefit from lesser amount of parameters or improved accuracy and exhibit systematic correspondence in representational complexity with the specified task. In contrast to conventional design patterns with a typical monotonic increase in the amount of features with increased depth, we observe that CNNs perform better when there is more learnable parameters in intermediate, with falloffs to earlier and later layers.

Via

Access Paper or Ask Questions