Literature DB >> 14587621

Speech segregation based on sound localization.

Nicoleta Roman1, DeLiang Wang, Guy J Brown.   

Abstract

At a cocktail party, one can selectively attend to a single voice and filter out all the other acoustical interferences. How to simulate this perceptual ability remains a great challenge. This paper describes a novel, supervised learning approach to speech segregation, in which a target speech signal is separated from interfering sounds using spatial localization cues: interaural time differences (ITD) and interaural intensity differences (IID). Motivated by the auditory masking effect, the notion of an "ideal" time-frequency binary mask is suggested, which selects the target if it is stronger than the interference in a local time-frequency (T-F) unit. It is observed that within a narrow frequency band, modifications to the relative strength of the target source with respect to the interference trigger systematic changes for estimated ITD and IID. For a given spatial configuration, this interaction produces characteristic clustering in the binaural feature space. Consequently, pattern classification is performed in order to estimate ideal binary masks. A systematic evaluation in terms of signal-to-noise ratio as well as automatic speech recognition performance shows that the resulting system produces masks very close to ideal binary ones. A quantitative comparison shows that the model yields significant improvement in performance over an existing approach. Furthermore, under certain conditions the model produces large speech intelligibility improvements with normal listeners.

Entities:  

Mesh:

Year:  2003        PMID: 14587621     DOI: 10.1121/1.1610463

Source DB:  PubMed          Journal:  J Acoust Soc Am        ISSN: 0001-4966            Impact factor:   1.840


  26 in total

Review 1.  How the owl tracks its prey--II.

Authors:  Terry T Takahashi
Journal:  J Exp Biol       Date:  2010-10-15       Impact factor: 3.312

2.  Determination of the potential benefit of time-frequency gain manipulation.

Authors:  Michael C Anzalone; Lauren Calandruccio; Karen A Doherty; Laurel H Carney
Journal:  Ear Hear       Date:  2006-10       Impact factor: 3.570

3.  Factors influencing glimpsing of speech in noise.

Authors:  Ning Li; Philipos C Loizou
Journal:  J Acoust Soc Am       Date:  2007-08       Impact factor: 1.840

4.  Use of a glimpsing model to understand the performance of listeners with and without hearing loss in spatialized speech mixtures.

Authors:  Virginia Best; Christine R Mason; Jayaganesh Swaminathan; Elin Roverud; Gerald Kidd
Journal:  J Acoust Soc Am       Date:  2017-01       Impact factor: 1.840

5.  Factors influencing intelligibility of ideal binary-masked speech: implications for noise reduction.

Authors:  Ning Li; Philipos C Loizou
Journal:  J Acoust Soc Am       Date:  2008-03       Impact factor: 1.840

Review 6.  Time-frequency masking for speech separation and its potential for hearing aid design.

Authors: 
Journal:  Trends Amplif       Date:  2008-10-30

7.  Effect of spectral resolution on the intelligibility of ideal binary masked speech.

Authors:  Ning Li; Philipos C Loizou
Journal:  J Acoust Soc Am       Date:  2008-04       Impact factor: 1.840

8.  The importance of processing resolution in "ideal time-frequency segregation" of masked speech and the implications for predicting speech intelligibility.

Authors:  Christopher Conroy; Virginia Best; Todd R Jennings; Gerald Kidd
Journal:  J Acoust Soc Am       Date:  2020-03       Impact factor: 1.840

9.  Comparison of a target-equalization-cancellation approach and a localization approach to source separation.

Authors:  Jing Mi; Matti Groll; H Steven Colburn
Journal:  J Acoust Soc Am       Date:  2017-11       Impact factor: 1.840

10.  Behavioral sensitivity to broadband binaural localization cues in the ferret.

Authors:  Peter Keating; Fernando R Nodal; Kohilan Gananandan; Andreas L Schulz; Andrew J King
Journal:  J Assoc Res Otolaryngol       Date:  2013-04-25
View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.