Skip to content

Opening book details…

Can I read 2020 International Joint Conference on Neural Networks (IJCNN) - Multi-level Visual Fusion Networks for Image Captioning on EtoBox?

2020 International Joint Conference on Neural Networks (IJCNN) - Multi-level Visual Fusion Networks for Image Captioning by Zhou, Dongming; Zhang, Canlong; Li, Zhixin; Wang, Zhiwen is a scholarly article available to read on EtoBox.

What is 2020 International Joint Conference on Neural Networks (IJCNN) - Multi-level Visual Fusion Networks for Image Captioning about?

Image captioning is a multi-modal complex task in machine learning. Traditional methods focus only on entities in visual strategy networks, and can't reason about the relationship between entities and attributes. There are problems of exposure bias and error accumulation in language strategy networks. To this end, this paper proposes a multi-level visual fusion network model based on reinforcement learning. In the visual strategy network, multi-level neural network modules are used to transform visual features into feature sets of visual knowledge. The fusion network generates function words that make the description more fluent, and is used for the interaction between the visual strategy network and the language strategy network. The self-criticism strategy gradient algorithm based on reinforcement learning in language strategy networks is used to achieve end-to-end optimization of visual fusion networks. We evaluated our model on the Flickr 30K and MS-COCO datasets, and verified the accuracy of the model and the diversity of model learning subtitles through experiments.Our model achieves better performance over state-of-the-art methods.

Author
Zhou, Dongming; Zhang, Canlong; Li, Zhixin; Wang, Zhiwen
Publisher
IEEE
Published
2020
Language
EN

More by Zhou, Dongming; Zhang, Canlong; Li, Zhixin; Wang, Zhiwen

Browse all works by Zhou, Dongming; Zhang, Canlong; Li, Zhixin; Wang, Zhiwen