About this document
Deep Learning for Language Identification by Maged Hamouda is a document available to read on EtoBox.
The document summarizes research on using deep learning for spoken language identification from audio clips. A deep neural network with convolutional and pooling layers (CNN-TDNN) is implemented to automatically learn features from spectrograms of speech samples, without relying on hand-coded features. The network is evaluated on two datasets of English, French, and German speech clips. On both datasets, the deep network achieves substantially higher identification accuracy compared to a shallow network, es
- Author
- Maged Hamouda
- Language
- EN