Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Martin Holeňa

On Difficulties of Attention Factorization through Shared Memory

Mar 31, 2024

Uladzislau Yorsh, Martin Holeňa, Ondřej Bojar, David Herel

Abstract:Transformers have revolutionized deep learning in numerous fields, including natural language processing, computer vision, and audio processing. Their strength lies in their attention mechanism, which allows for the discovering of complex input relationships. However, this mechanism's quadratic time and memory complexity pose challenges for larger inputs. Researchers are now investigating models like Linear Unified Nested Attention (Luna) or Memory Augmented Transformer, which leverage external learnable memory to either reduce the attention computation complexity down to linear, or to propagate information between chunks in chunk-wise processing. Our findings challenge the conventional thinking on these models, revealing that interfacing with the memory directly through an attention operation is suboptimal, and that the performance may be considerably improved by filtering the input signal before communicating with memory.

* 2 pages of main content, 8 pages in total, published as a Tiny Paper at ICLR 2024

Via

Access Paper or Ask Questions

Video Scene Location Recognition with Neural Networks

Sep 21, 2023

Lukáš Korel, Petr Pulc, Jiří Tumpach, Martin Holeňa

Abstract:This paper provides an insight into the possibility of scene recognition from a video sequence with a small set of repeated shooting locations (such as in television series) using artificial neural networks. The basic idea of the presented approach is to select a set of frames from each scene, transform them by a pre-trained singleimage pre-processing convolutional network, and classify the scene location with subsequent layers of the neural network. The considered networks have been tested and compared on a dataset obtained from The Big Bang Theory television series. We have investigated different neural network layers to combine individual frames, particularly AveragePooling, MaxPooling, Product, Flatten, LSTM, and Bidirectional LSTM layers. We have observed that only some of the approaches are suitable for the task at hand.

Via

Access Paper or Ask Questions

Using Artificial Neural Networks to Determine Ontologies Most Relevant to Scientific Texts

Sep 17, 2023

Lukáš Korel, Alexander S. Behr, Norbert Kockmann, Martin Holeňa

Figure 1 for Using Artificial Neural Networks to Determine Ontologies Most Relevant to Scientific Texts

Figure 2 for Using Artificial Neural Networks to Determine Ontologies Most Relevant to Scientific Texts

Figure 3 for Using Artificial Neural Networks to Determine Ontologies Most Relevant to Scientific Texts

Figure 4 for Using Artificial Neural Networks to Determine Ontologies Most Relevant to Scientific Texts

Abstract:This paper provides an insight into the possibility of how to find ontologies most relevant to scientific texts using artificial neural networks. The basic idea of the presented approach is to select a representative paragraph from a source text file, embed it to a vector space by a pre-trained fine-tuned transformer, and classify the embedded vector according to its relevance to a target ontology. We have considered different classifiers to categorize the output from the transformer, in particular random forest, support vector machine, multilayer perceptron, k-nearest neighbors, and Gaussian process classifiers. Their suitability has been evaluated in a use case with ontologies and scientific texts concerning catalysis research. From results we can say the worst results have random forest. The best results in this task brought support vector machine classifier.

Via

Access Paper or Ask Questions

Two Gaussian Approaches to Black-Box Optomization

Nov 28, 2014

Lukáš Bajer, Martin Holeňa

Abstract:Outline of several strategies for using Gaussian processes as surrogate models for the covariance matrix adaptation evolution strategy (CMA-ES).

* 9 pages

Via

Access Paper or Ask Questions