Opening book details…
Can I read Multi-Modal Dense Video Captioning Framework on EtoBox?
Multi-Modal Dense Video Captioning Framework by nanavathakashchandra is a document available to read on EtoBox.
What is Multi-Modal Dense Video Captioning Framework about?
The document outlines a project focused on developing a multi-modal dense video captioning framework that generates accurate captions for events in untrimmed videos by integrating audio, visual, and speech data. It highlights the challenges of combining different data types and aims to enhance the quality of representations using advanced architectures like AST, BERT, and ViViT. The project will be evaluated using the ActivityNet Captions dataset to establish benchmark performance.
- Author
- nanavathakashchandra
- Language
- EN