Literature DB >> 33744714

Speaker recognition based on deep learning: An overview.

Zhongxin Bai1, Xiao-Lei Zhang2.   

Abstract

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarization. Finally, we survey robust speaker recognition from the perspectives of domain adaptation and speech enhancement, which are two major approaches of dealing with domain mismatch and noise problems. Popular and recently released corpora are listed at the end of the paper.
Copyright © 2021 Elsevier Ltd. All rights reserved.

Entities:  

Keywords:  Deep learning; Robust speaker recognition; Speaker diarization; Speaker identification; Speaker recognition; Speaker verification

Year:  2021        PMID: 33744714     DOI: 10.1016/j.neunet.2021.03.004

Source DB:  PubMed          Journal:  Neural Netw        ISSN: 0893-6080


  4 in total

1.  Deep Learning-Based Device-Free Localization Scheme for Simultaneous Estimation of Indoor Location and Posture Using FMCW Radars.

Authors:  Jeongpyo Lee; Kyungeun Park; Youngok Kim
Journal:  Sensors (Basel)       Date:  2022-06-12       Impact factor: 3.847

2.  A Novel Framework for Open-Set Authentication of Internet of Things Using Limited Devices.

Authors:  Keju Huang; Junan Yang; Pengjiang Hu; Hui Liu
Journal:  Sensors (Basel)       Date:  2022-03-30       Impact factor: 3.576

3.  Frequency, Time, Representation and Modeling Aspects for Major Speech and Audio Processing Applications.

Authors:  Juraj Kacur; Boris Puterka; Jarmila Pavlovicova; Milos Oravec
Journal:  Sensors (Basel)       Date:  2022-08-22       Impact factor: 3.847

4.  Identifying the Strength Level of Objects' Tactile Attributes Using a Multi-Scale Convolutional Neural Network.

Authors:  Peng Zhang; Guoqi Yu; Dongri Shan; Zhenxue Chen; Xiaofang Wang
Journal:  Sensors (Basel)       Date:  2022-03-01       Impact factor: 3.576

  4 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.