Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering.

Literature DB >> 29993847

Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering.

Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, Dacheng Tao.

Abstract

Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.

Year: 2018 PMID： 29993847 DOI： 10.1109/TNNLS.2018.2817340

Source DB: PubMed Journal: IEEE Trans Neural Netw Learn Syst ISSN： 2162-237X Impact factor: 10.451

Keyword Cloud
Cited

7 in total

1. Bilinear pooling in video-QA: empirical challenges and motivational drift from neurological parallels.

Authors: Thomas Winterbottom; Sarah Xiao; Alistair McLean; Noura Al Moubayed
Journal: PeerJ Comput Sci Date: 2022-06-03

2. Multi-path x-D Recurrent Neural Networks for Collaborative Image Classification.

Authors: Riqiang Gao; Yuankai Huo; Shunxing Bao; Yucheng Tang; Sanja L Antic; Emily S Epstein; Steve Deppen; Alexis B Paulson; Kim L Sandler; Pierre P Massion; Bennett A Landman
Journal: Neurocomputing Date: 2020-02-15 Impact factor: 5.719

3. Toward Accurate Visual Reasoning With Dual-Path Neural Module Networks.

Authors: Ke Su; Hang Su; Jianguo Li; Jun Zhu
Journal: Front Robot AI Date: 2020-08-21

4. Multi-Modal Explicit Sparse Attention Networks for Visual Question Answering.

Authors: Zihan Guo; Dezhi Han
Journal: Sensors (Basel) Date: 2020-11-26 Impact factor: 3.576

5. Deep transfer learning of structural magnetic resonance imaging fused with blood parameters improves brain age prediction.

Authors: Bingyu Ren; Yingtong Wu; Liumei Huang; Zhiguo Zhang; Bingsheng Huang; Huajie Zhang; Jinting Ma; Bing Li; Xukun Liu; Guangyao Wu; Jian Zhang; Liming Shen; Qiong Liu; Jiazuan Ni
Journal: Hum Brain Mapp Date: 2021-12-16 Impact factor: 5.038

6. Deep Modular Bilinear Attention Network for Visual Question Answering.

Authors: Feng Yan; Wushouer Silamu; Yanbing Li
Journal: Sensors (Basel) Date: 2022-01-28 Impact factor: 3.576

7. Interpretable disease prediction using heterogeneous patient records with self-attentive fusion encoder.

Authors: Heeyoung Kwak; Jooyoung Chang; Byeongjin Choe; Sangmin Park; Kyomin Jung
Journal: J Am Med Inform Assoc Date: 2021-09-18 Impact factor: 7.942

7 in total