Literature DB >> 33578956

Image Captioning Using Motion-CNN with Object Detection.

Kiyohiko Iwamura1, Jun Younes Louhi Kasahara1, Alessandro Moro1,2, Atsushi Yamashita1, Hajime Asama1.   

Abstract

Automatic image captioning has many important applications, such as the depiction of visual contents for visually impaired people or the indexing of images on the internet. Recently, deep learning-based image captioning models have been researched extensively. For caption generation, they learn the relation between image features and words included in the captions. However, image features might not be relevant for certain words such as verbs. Therefore, our earlier reported method included the use of motion features along with image features for generating captions including verbs. However, all the motion features were used. Since not all motion features contributed positively to the captioning process, unnecessary motion features decreased the captioning accuracy. As described herein, we use experiments with motion features for thorough analysis of the reasons for the decline in accuracy. We propose a novel, end-to-end trainable method for image caption generation that alleviates the decreased accuracy of caption generation. Our proposed model was evaluated using three datasets: MSR-VTT2016-Image, MSCOCO, and several copyright-free images. Results demonstrate that our proposed method improves caption generation performance.

Entities:  

Keywords:  deep learning; image captioning; motion estimation; object detection

Mesh:

Year:  2021        PMID: 33578956      PMCID: PMC7916682          DOI: 10.3390/s21041270

Source DB:  PubMed          Journal:  Sensors (Basel)        ISSN: 1424-8220            Impact factor:   3.576


  1 in total

1.  A Lightweight Optical Flow CNN -Revisiting Data Fidelity and Regularization.

Authors:  Tak-Wai Hui; Xiaoou Tang; Chen Change Loy
Journal:  IEEE Trans Pattern Anal Mach Intell       Date:  2021-07-01       Impact factor: 6.226

  1 in total
  1 in total

1.  An accurate generation of image captions for blind people using extended convolutional atom neural network.

Authors:  Tejal Tiwary; Rajendra Prasad Mahapatra
Journal:  Multimed Tools Appl       Date:  2022-07-15       Impact factor: 2.577

  1 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.