Can I read Modality Fusion for Audio-Visual ID on EtoBox?
Modality Fusion for Audio-Visual ID by dxdmlq is a document available to read on EtoBox.
What is Modality Fusion for Audio-Visual ID about?
This paper presents a comparative analysis of three modality fusion strategies for audio-visual person identification and verification using voice and face data. The study employs deep learning techniques, specifically a one-dimensional convolutional neural network for voice and a pre-trained VGGFace2 network for face recognition, achieving an accuracy of 98.37% in person identification tasks. The results indicate that feature fusion of gammatonegram and facial features outperforms other methods, while the
- Author
- dxdmlq
- Language
- EN