About this document
CNN-LSTM Image Captioning Model by Tsegazewold Kinfu is a document available to read on EtoBox.
The document discusses image captioning using a CNN-LSTM architecture. It involves using a CNN to extract visual features from images, and an LSTM to generate natural language captions describing the images. Specifically, it uses a pre-trained ResNet50 CNN to extract features from images, and an LSTM language model to translate those features into English captions. It provides details on how the CNN-LSTM model is built and trained on a dataset containing images and corresponding text descriptions, to learn
- Author
- Tsegazewold Kinfu
- Language
- EN