Loading...

Audio-visual Speaker Verification Using Identity Based Articulatory Features

Keshavarz Chafjiri, Amir Hossein | 2020

669 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 53597 (05)
  4. University: Sharif University of Technology
  5. Department: Electrical Engineering
  6. Advisor(s): Ghaem Maghami, Shahrokh
  7. Abstract:
  8. Many companies feel the need of automatic person identification for tracking entrance and exit of their employers. This system will help the company in reducing security costs, waste of time and increasing security. In this thesis a novel idea for 3d convolutional neural network text independent speaker verification system will introduced. Since many scenarios visionary data like lip motion, besides speech data, is available, in this thesis both of these data streams are used for developing designed speaker verification system. Multimodalily helps the verification system to be more robust to the noise and attacks. One of the main challenges in designing multimodal neural network speaker verification systems is to use an architecture that discriminates speakers the most. So, we have used Siamese architecture for discriminating speakers. In this thesis mfec features for speech and raw lip images are used as inputs of neural networks.The proposed method is evaluated with VoxCeleb dataset and EER metrics. It is shown that proposed method reaches 4.5% EER that is better performance than i-vector and similar articles
  9. Keywords:
  10. Speaker Verification ; Equal Error Rate (EER) ; Siamese Architecture ; VoxCeleb Dataset ; Three Dimentional Convolutional Neural Networks

 Digital Object List

 Bookmark

No TOC