Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Shuki Shimizu

Joint learning of images and videos with a single Vision Transformer

Aug 21, 2023

Shuki Shimizu, Toru Tamaki

Abstract:In this study, we propose a method for jointly learning of images and videos using a single model. In general, images and videos are often trained by separate models. We propose in this paper a method that takes a batch of images as input to Vision Transformer IV-ViT, and also a set of video frames with temporal aggregation by late fusion. Experimental results on two image datasets and two action recognition datasets are presented.

* MVA2023 (18th International Conference on Machine Vision Applications), Hamamatsu, Japan, 23-25 July 2023

Via

Access Paper or Ask Questions