Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 ARTSTREAM: a neural network model of auditory scene analysis and source segregation.

Literature DB >> 15109681

ARTSTREAM: a neural network model of auditory scene analysis and source segregation.

Stephen Grossberg¹, Krishna K Govindarajan, Lonce L Wyse, Michael A Cohen.

Abstract

Multiple sound sources often contain harmonics that overlap and may be degraded by environmental noise. The auditory system is capable of teasing apart these sources into distinct mental objects, or streams. Such an 'auditory scene analysis' enables the brain to solve the cocktail party problem. A neural network model of auditory scene analysis, called the ARTSTREAM model, is presented to propose how the brain accomplishes this feat. The model clarifies how the frequency components that correspond to a given acoustic source may be coherently grouped together into a distinct stream based on pitch and spatial location cues. The model also clarifies how multiple streams may be distinguished and separated by the brain. Streams are formed as spectral-pitch resonances that emerge through feedback interactions between frequency-specific spectral representations of a sound source and its pitch. First, the model transforms a sound into a spatial pattern of frequency-specific activation across a spectral stream layer. The sound has multiple parallel representations at this layer. A sound's spectral representation activates a bottom-up filter that is sensitive to the harmonics of the sound's pitch. This filter activates a pitch category which, in turn, activates a top-down expectation that is also sensitive to the harmonics of the pitch. Resonance develops when the spectral and pitch representations mutually reinforce one another. Resonance provides the coherence that allows one voice or instrument to be tracked through a noisy multiple source environment. Spectral components are suppressed if they do not match harmonics of the top-down expectation that is read-out by the selected pitch, thereby allowing another stream to capture these components, as in the 'old-plus-new heuristic' of Bregman. Multiple simultaneously occurring spectral-pitch resonances can hereby emerge. These resonance and matching mechanisms are specialized versions of Adaptive Resonance Theory, or ART, which clarifies how pitch representations can self-organize during learning of harmonic bottom-up filters and top-down expectations. The model also clarifies how spatial location cues can help to disambiguate two sources with similar spectral cues. Data are simulated from psychophysical grouping experiments, such as how a tone sweeping upwards in frequency creates a bounce percept by grouping with a downward sweeping tone due to proximity in frequency, even if noise replaces the tones at their intersection point. Illusory auditory percepts are also simulated, such as the auditory continuity illusion of a tone continuing through a noise burst even if the tone is not present during the noise, and the scale illusion of Deutsch whereby downward and upward scales presented alternately to the two ears are regrouped based on frequency proximity, leading to a bounce percept. Since related sorts of resonances have been used to quantitatively simulate psychophysical data about speech perception, the model strengthens the hypothesis that ART-like mechanisms are used at multiple levels of the auditory system. Proposals for developing the model to explain more complex streaming data are also provided.

Mesh：

Year: 2004 PMID： 15109681 DOI： 10.1016/j.neunet.2003.10.002

Source DB: PubMed Journal: Neural Netw ISSN： 0893-6080

Keyword Cloud
Cited

17 in total

1. A cocktail party with a cortical twist: how cortical mechanisms contribute to sound segregation.

Authors: Mounya Elhilali; Shihab A Shamma
Journal: J Acoust Soc Am Date: 2008-12 Impact factor: 1.840

2. Segregated audio-tactile events destabilize the bimanual coordination of distinct rhythms.

Authors: Julien Lagarde; Gregory Zelic; Denis Mottet
Journal: Exp Brain Res Date: 2012-05-09 Impact factor: 1.972

3. Bayesian inference in auditory scenes.

Authors: Mounya Elhilali
Journal: Conf Proc IEEE Eng Med Biol Soc Date: 2013

Review 4. Cortical and subcortical predictive dynamics and learning during perception, cognition, emotion and action.

Authors: Stephen Grossberg
Journal: Philos Trans R Soc Lond B Biol Sci Date: 2009-05-12 Impact factor: 6.237

5. Selective entrainment of brain oscillations drives auditory perceptual organization.

Authors: Jordi Costa-Faidella; Elyse S Sussman; Carles Escera
Journal: Neuroimage Date: 2017-07-27 Impact factor: 6.556

6. Retrograde adaptive resonance theory based on the role of nitric oxide in long-term potentiation.

Authors: Peng Jia; Junsong Yin; Dewen Hu; Zongtan Zhou
Journal: J Comput Neurosci Date: 2007-04-03 Impact factor: 1.621

7. Hearing an illusory vowel in noise: suppression of auditory cortical activity.

Authors: Lars Riecke; Mieke Vanbussel; Lars Hausfeld; Deniz Başkent; Elia Formisano; Fabrizio Esposito
Journal: J Neurosci Date: 2012-06-06 Impact factor: 6.167

8. Individual differences in sound-in-noise perception are related to the strength of short-latency neural responses to noise.

Authors: Ekaterina Vinnik; Pavel M Itskov; Evan Balaban
Journal: PLoS One Date: 2011-02-28 Impact factor: 3.240

9. An adaptation level theory of tinnitus audibility.

Authors: Grant D Searchfield; Kei Kobayashi; Michael Sanders
Journal: Front Syst Neurosci Date: 2012-06-13

10. Modelling the emergence and dynamics of perceptual organisation in auditory streaming.

Authors: Robert W Mill; Tamás M Bőhm; Alexandra Bendixen; István Winkler; Susan L Denham
Journal: PLoS Comput Biol Date: 2013-03-14 Impact factor: 4.475