<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JSIP</journal-id><journal-title-group><journal-title>Journal of Signal and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2159-4465</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jsip.2012.32032</article-id><article-id pub-id-type="publisher-id">JSIP-19576</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  About Multichannel Speech Signal Extraction and Separation Techniques
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>del</surname><given-names>Hidri</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Souad</surname><given-names>Meddeb</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hamid</surname><given-names>Amiri</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Signal Image and Technology of Information Laboratory, National Engineering School of Tunis, Tunis, Tunisia</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>hidri_adel@yahoo.fr(DH)</email>;<email>mmemeddeb@gmail.com(SM)</email>;<email>hamidlamiri@yahoo.com(HA)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>30</day><month>05</month><year>2012</year></pub-date><volume>03</volume><issue>02</issue><fpage>238</fpage><lpage>247</lpage><history><date date-type="received"><day>March</day>	<month>14th,</month>	<year>2012</year></date><date date-type="rev-recd"><day>April</day>	<month>12th,</month>	<year>2012</year>	</date><date date-type="accepted"><day>May</day>	<month>8th,</month>	<year>2012</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The extraction of a desired speech signal from a noisy environment has become a challenging issue. In the recent years, the scientific community has particularly focused on multichannel techniques which are dealt with in this review. In fact, this study tries to classify these multichannel techniques into three main ones: Beamforming, Independent Component Analysis (ICA) and Time Frequency (T-F) masking. This paper also highlights their advantages and drawbacks. However these previously mentioned techniques could not afford satisfactory results. This fact leads to the idea that a combination of those techniques, which is depicted along this study, may probably provide more efficient results. Indeed, giving the fact that those approaches are still be considered as being not totally efficient, has led us to review these mentioned above in the hope that further researches will provide this domain with suitable innovations.
 
</p></abstract><kwd-group><kwd>Beamforming; ICA; T-F Masking; BSS; Multichannel; Speech Separation; Microphone Array</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Most audio signals result from the mixing of several sound sources. In many applications, there is a need to separate the multiple sources or extract a source of interest while reducing undesired interfering signals and noise. The estimated signals may then be either directly listened to or further processed, giving rise to a wide range of applications such as hearing aids, human computer interaction, surveillance, and hands-free telephony [<xref ref-type="bibr" rid="scirp.19576-ref1">1</xref>].</p><p>The extraction of a desired speech signal from a mixture of multiple signals is classically referred to as the “cocktail party problem” [2,3], where different conversations occur simultaneously and independently of each other.</p><p>The human auditory system shows a remarkable ability to segregate only one conversation in a highly noisy environment, such as in a cocktail party environment. However, it remains extremely challenging for machines to replicate even part of such functionalities. Despite being studied for decades, the cocktail party problem remains a scientific challenge that demands further research efforts [<xref ref-type="bibr" rid="scirp.19576-ref4">4</xref>].</p><p>As highlighted in some recent works [<xref ref-type="bibr" rid="scirp.19576-ref5">5</xref>], using a single channel is not possible to improve both intelligibility and quality of the recovered signal at the same time. Quality can be improved at the expense of sacrificing intelligibility. A way to overcome this limitation is to add some spatial information to the time/frequency information available in the single channel case. Actually, this additional information could be obtained by using two or more channel of noisy speech named multichannel.</p><p>Three techniques of Multi Channel Speech Signal Separation and Extraction (MCSSE) can be defined. The first two techniques are designed to determined and overdetermined mixtures (when the number of sources is smaller than or equal to the number of mixtures) and the third is designed to underdetermined mixtures (when the number of sources is larger than the number of mixtures). The former is based on two famous approaches, the Blind Source Separation (BSS) techniques [5-7] and the Beamforming techniques [8-10].</p><p>BSS aims at separating all the involved sources, by exploiting their independent statistical properties, regardless their attribution to the desired or interfering sources.</p><p>On the other hand, the Beamforming techniques, concentrate on enhancing the sum of the desired sources while treating all other signals as interfering sources. While the latter uses the knowledge of speech signal properties for separation.</p><p>One popular approach to sparsity based separation is T-F masking [11-13]. This approach is a special case of non-linear time-varying filtering that estimates the desired source from a mixture signal by applying a T-F mask that attenuates T-F points associated with interfering signals while preserving T-F points where the signal of interest is dominant.</p><p>In the last years, the researches in this area based their approaches on combination techniques as ICA and binary T-F masking [<xref ref-type="bibr" rid="scirp.19576-ref14">14</xref>], Beamforming and a time frequency binary mask [<xref ref-type="bibr" rid="scirp.19576-ref15">15</xref>].</p><p>This paper is concerned with a survey of the main ideas in the area of speech separation and extraction from a multiple microphones.</p><p>The following sections of this paper are organized as follows: in Section 2, the problem of speech separation and extraction is formulated. In Section 3, we describe some of the most techniques which have been used in MCSSE systems, such as Beamforming, ICA and T-F masking techniques. Section 4 brings to the surface the most recent methods for MCSSE systems, where combined techniques, seen previously, are used. In Section 5, the presented methods will be discussed by giving some of their advantages and limits. Finally, Section 6 gives a synopsis of the whole paper and conveys some futures works.</p></sec><sec id="s2"><title>2. Problem Formulation</title><p>There are many scenarios where audio mixtures can be obtained. This results in different characteristics of the sources and the mixing process that can be exploited by the separation methods. The observed spatial properties of audio signals depend on the spatial distribution of a sound source, the sound scene acoustics, the distance between the source and the microphones, and the directivity of the microphones.</p><p>In general, the problem of MCSSE is stated to be the process of estimating the signals from N unobserved sources, given from M microphones, which arises when the signals from the N unobserved sources are linearly mixed together as presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><p>The signal recorded at the j<sup>th</sup> microphone can be modeled as:</p><disp-formula id="scirp.19576-formula38124"><label>(1)</label><graphic position="anchor" xlink:href="15-3400188\0ddc4931-7df6-4f52-855b-7da22a2ad5ed.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="15-3400188\34ab49c9-c746-4ac2-8452-7f582dbe8a89.jpg" /> and <img src="15-3400188\ac165a8c-1cbf-48c9-8b4f-183db8778d65.jpg" /> are the source and mixture signals respectively, h<sub>ji</sub> is a P-point Room Impulse Response (RIR) from source i to microphone j, P is the number of paths between each source-microphone pair and <img src="15-3400188\a01f6a5c-41b0-40d5-8896-431bb59e7b63.jpg" /> is the delay of the p<sup>th</sup> path from source j to microphone i [9-14]. This model is the most natural mixing model, encountered in live recordings called echoic mixtures.</p><p>In free-reverberation environments&#160;(p = 1), the samples of each source signal can arrive at the microphones only from the line of sight path, and the attenuation and delay of source i would be determined by the physical position of the source relative to the microphones. This model, called anechoic mixing, is described by the following equation obtained from the previous equation:</p><disp-formula id="scirp.19576-formula38125"><label>(2)</label><graphic position="anchor" xlink:href="15-3400188\0ff51206-6ac6-4242-9aa5-fac5c583cae8.jpg"  xlink:type="simple"/></disp-formula><p>The instantaneous mixing model is a specific case of the anechoic mixing model where the samples of each source arrive at the microphones at the same time <img src="15-3400188\ad1efcb2-b74a-4c1a-8479-b9b134a96fba.jpg" /> with differing attenuations, each element of the mixing matrix <img src="15-3400188\f9573938-567b-400f-98f2-317fedeb7d04.jpg" /> is a scalar that represents the amplitude scaling between source i and microphone j. From the Equation (2), instantaneous mixing model can be expressed as:</p><disp-formula id="scirp.19576-formula38126"><label>(3)</label><graphic position="anchor" xlink:href="15-3400188\8b2675c8-f167-4e91-bf8e-2a344a7e9538.jpg"  xlink:type="simple"/></disp-formula></sec><sec id="s3"><title>3. MCSSE Techniques</title><sec id="s3_1"><title>3.1. Beamforming Technique</title><p>Beamforming is a class of algorithms for multichannel signal processing. The term Beamforming refers to the design of a spatio-temporal filter which operates on the outputs of the microphone array [<xref ref-type="bibr" rid="scirp.19576-ref8">8</xref>]. This spatial filter can be expressed in terms of dependence upon angle and frequency. Beamforming is accomplished by filtering the microphone signals and combining the outputs to extract (by constructive combining) the desired signal and reject (by destructive combining) interfering signals according to their spatial location [<xref ref-type="bibr" rid="scirp.19576-ref9">9</xref>].</p><p>Beamforming for broadband signals like speech can, in general, be performed in the time domain or frequency domain. In time domain Beamforming, a Finite Impulse Response (FIR) filter is applied to each microphone signal, and the filter outputs combined to form the Beamformer output. Beamforming can be performed by computing multichannel filters whose output is <img src="15-3400188\518e65b0-6824-43fa-b22d-baaaefa485e7.jpg" /> an estimate of the desired source signal as shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>The output can be expressed as:</p><disp-formula id="scirp.19576-formula38127"><label>(4)</label><graphic position="anchor" xlink:href="15-3400188\bdc5c267-1995-49c0-8c05-36aacefc9dec.jpg"  xlink:type="simple"/></disp-formula><p>where P – 1 is the number of delays in each of the N filters.</p><p>In frequency domain Beamforming, the microphone signal is separated into narrowband frequency bins using a Short-Time Fourier Transform (STFT), and the data in each frequency bin is processed separately.</p><p>Beamforming techniques can be broadly classiﬁed as being either data-independent or data-dependent. Data independent or deterministic Beamformers are so named because their filters do not depend on the microphone signals and are chosen to approximate a desired response. Conversely, data-dependent or statistically optimum Beamforming techniques are been so called because their filters are based on the statistics of the arriving data to optimize some function that makes the Beamformer optimum in some sense.</p><sec id="s3_1_1"><title>3.1.1. Deterministic Beamformer</title><p>The filters in a deterministic Beamformer do not depend on the microphone signals and are chosen to approximate a desired response. For example, we may wish to receive any signal arriving from a certain direction, in which case the desired response is unity over at that direction. As another example, we may know that there is interference operating at a certain frequency and arriving from a certain direction, in which case the desired response at that frequency and direction is zero. The simplest deterministic Beamforming technique is delay-and-sum Beamforming, where the signals at the microphones are delayed and then summed in order to combine the signal arriving from the direction of the desired source coherently, expecting that the interference components arriving from off the desired direction cancel to a certain extent by destructive combining. The delay-and-sum Beamformer as shown in <xref ref-type="fig" rid="fig3">Figure 3</xref> is simple in its implementation and provides easy steering of the beam towards the desired source. Assuming that the broadband signal can be decomposed into narrowband frequency bins, the delays can be approximated by phase shifts in each frequency band.</p><p>The performance of the delay-and-sum Beamformer in reverberant environments is often insufficient. A more general processing model is the filter-and-sum Beamformer as shown in <xref ref-type="fig" rid="fig4">Figure 4</xref> where, before summation, each microphone signal is filtered with FIR filters of order M. This structure, designed for multipath environments namely reverberant enclosures, replaces the simpler delay compensator with a matched filter. It is one of the simplest Beamforming techniquesbut still gives a very good performance.</p><p>As it has been shown that the deterministic Beamformer is far from being fully manipulated independently from the microphone signals, the statistically optimal Beamformer is tightly linked and tied to the statistical properties of the received signals.</p></sec><sec id="s3_1_2"><title>3.1.2. Statistically Optimum Beamformer</title><p>Statistically optimal Beamformers are designed basing on the statistical properties of the desired and interference signals. In this category, the filters designs are based on the statistics of the arriving data to optimize some function that makes the Beamformer optimum in some sense. Several criteria can be applied in the design of the Beamformer, e.g., maximum signal-to-noise ratio (MSNR), minimum mean-squared error (MMSE), minimum variance distortionless response (MVDR) and linear constraint minimum variance (LCMV). A summary of several design criteria can be found in [<xref ref-type="bibr" rid="scirp.19576-ref10">10</xref>]. In general, they aim at enhancing the desired signals, while rejecting the interfering signals.</p><p><xref ref-type="fig" rid="fig5">Figure 5</xref> depicts the block diagram of Frost Beamformer or an adaptive filter-and-sum Beamformer as proposed in [<xref ref-type="bibr" rid="scirp.19576-ref16">16</xref>], where the filter coefficients are adapted using a constrained version of the Least Mean-Square (LMS) algorithm. The LMS is used to minimize the noise power at the output while maintaining a constraint</p></sec></sec></sec></body><back><ref-list><title>References</title><ref id="scirp.19576-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">M. Brandstein and D. Ward, “Microphone Arrays: Signal Processing Techniques and Applications,” Digital Signal Processing, 2001, Springer.</mixed-citation></ref><ref id="scirp.19576-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">C. Cherry: “Some Experiments on the Recognition of Speech, with One and with Two Ears,” Journal of the Acoustical Society of America, Vol. 25, No. 5, 1953, pp. 975–979. doi:10.1121/1.1907229</mixed-citation></ref><ref id="scirp.19576-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">S. Haykin and Z. Chen, “The Cocktail Party Problem,” Journals of Neural Computation, Vol. 17, No. 9, 2005, pp. 1875-1902. doi:10.1162/0899766054322964</mixed-citation></ref><ref id="scirp.19576-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">D. L. Wang and G. J. Brown, “Computational Auditory Scene Analysis: Principles Algorithms and Applications,” Wiley, New York, 2006. 10.1109/TNN.2007.913988</mixed-citation></ref><ref id="scirp.19576-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">J. Benesty, S. Makino and J. Chen, “Speech Enhancement,” Signal and Communication Technology, Springer, Berlin, 2005.</mixed-citation></ref><ref id="scirp.19576-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">S. Douglas and M. Gupta, “Convolutive Blind Source Separation for Audio Signals,” Blind Speech Separation, Springer, Berlin, 2007.  
doi:10.1007/978-1-4020-6479-1_1</mixed-citation></ref><ref id="scirp.19576-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">H. Sawada, S. Araki and S. Makino “Frequency-Domain Blind Source Separation,” Blind Speech Separation, Springer, Berlin, 2007. doi:10.1007/3-540-27489-8_13</mixed-citation></ref><ref id="scirp.19576-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">S. Markovich, S. Gannot and I. Cohen, “Multichannel Eigen Space Beamforming in a Reverberant Noisy Environment with Multiple Interfering Speech Signals,” IEEE Transactions on Audio, Speech, and Language Processing, Vol. 17, No. 6, 2009, pp. 1071-1086. 
doi:10.1109/TASL.2009.2016395</mixed-citation></ref><ref id="scirp.19576-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">M. A. Dmour and M. Davies “A New Framework for Underdetermined Speech Extraction Using Mixture of Beamformers,” IEEE Transactions on Audio, Speech, and Language Processing, Vol. 19, No. 3, 2011, pp. 445-457.  
doi:10.1109/TASL.2010.2049514</mixed-citation></ref><ref id="scirp.19576-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">J. Benesty, J. Chen and Y. Huang, “Conventional Beamforming Techniques,” Microphone Array Signal Processing, Springer, Berlin, 2008. doi:10.1121/1.3124775</mixed-citation></ref><ref id="scirp.19576-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">V. G. Reju, S. N. Koh and I. Y. Soon, “Underdetermined Convolutive Blind Source Separation via Time-Frequency Masking,” IEEE Transactions on Audio, Speech, and Language Processing, Vol. 18, No. 1, 2010, pp. 101-116.  
doi:10.1109/TASL.2009.2024380</mixed-citation></ref><ref id="scirp.19576-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">O. Yilmaz and S. Rickard, “Blind Separation of Speech Mixtures via Time-Frequency Masking,” IEEE Transactions on Signal Processing, Vol. 52, 2004, pp. 1830-1847.  
doi:10.1109/TSP.2004.828896</mixed-citation></ref><ref id="scirp.19576-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">J. Freudenberger and S. Stenzel, “Time-Frequency Masking for Convolutive and Noisy Mixtures,” Workshop on Hands-Free Speech Communication and Microphone Arrays, 2011, pp. 104-108.  
doi:10.1109/HSCMA.2011.5942374</mixed-citation></ref><ref id="scirp.19576-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">T. Jan, W. Wang and D. L. Wang, “A Multistage Approach to Blind Separation of Convolutive Speech Mixtures,” Speech Communication, Vol. 53, 2011, pp. 524-539. 
doi:10.1016/j.specom.2011.01.002</mixed-citation></ref><ref id="scirp.19576-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">J. Cermak, S. Araki, H. Sawada and S. Makino, “Blind Speech Separation by Combining Beamformers and a Time Frequency Binary Mask,” IEEE International Conference on Acoustics, Speech and Signal Processing, Honolulu, 2007, pp. I-145-I-148.  </mixed-citation></ref><ref id="scirp.19576-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">O. Frost, “An Algorithm for Linearly Constrained Adaptive Array Processing,” Proceedings of the IEEE, Vol. 60, No. 8, 1972, pp. 926-935.</mixed-citation></ref><ref id="scirp.19576-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">E. A. P. Habets, J. Benesty, I. Cohen, S. Gannot and J. Dmochowski, “New Insights into the MVDR Beamformer in Room Acoustics,” IEEE Transactions on Audio, Speech, and Language Processing, 2010, pp. 158-170. 
doi:10.1109/TASL.2009.2024731</mixed-citation></ref><ref id="scirp.19576-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">L. Griffiths and C. Jim, “An Alternative Approach to Linearly Constrained Adaptive Beamforming,” IEEE Transactions on Antennas and Propagation, Vol. 30, No. 1, 1982, pp. 27-34. doi:10.1109/TAP.1982.1142739</mixed-citation></ref><ref id="scirp.19576-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">S. Gannot and I. Cohen “Adaptive Beamforming and Post filtering,” Speech Processing, Springer, Berlin, 2007, pp. 199-228.</mixed-citation></ref><ref id="scirp.19576-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">A. Spriet, M. Moonen and J. Wouters, “Spatially Pre-Processed Speech Distortion Weighted Multi-Channel Wiener Filtering for Noise Reduction,” Signal Processing, Vol. 84, No. 12, 2004, pp. 2367-2387. 
doi:10.1016/j.sigpro.2004.07.028</mixed-citation></ref><ref id="scirp.19576-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">P. Comon, “Independent Component Analysis, a New Concept,” Signal Processing, Vol. 36, No, 3, 1994, pp. 287-314. doi:10.1016/0165-1684(94)90029-9</mixed-citation></ref><ref id="scirp.19576-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Z. Koldovsky and P. Tichavsky, “Time-Domain Blind Audio Source Separation Using Advanced ICA Methods,” Interspeech, Antwerp Belgium, 2007, pp. 846-849.</mixed-citation></ref><ref id="scirp.19576-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">S. Makino, H. Sawada, R. Mukai and S. Araki, “Blind Source Separation of Convolutive Mixtures of Speech in Frequency Domain,” IEICE Transactions on Fundamentals of Electronics Communications and Computer Sciences, No. 7, 2005, pp. 1640-1655.  
doi:10.1093/ietfec/e88-a.7.1640</mixed-citation></ref><ref id="scirp.19576-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">A. Sarmiento, I. Durán-Díaz, S. Cruces and P. Aguilera, “Generalized Method for Solving the Permutation Problem in Frequency-Domain Blind Source Separation of Convolved Speech Signals,” Interspeech, 2011, pp. 565-568.</mixed-citation></ref><ref id="scirp.19576-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">R. Mazur and A. Mertins, “A Sparsity Based Criterion for Solving the Permutation Ambiguity in Convolutive Blind Source Separation,” IEEE International Conference on Acoustics, Speech and Signal Processing, Prague Czech Republic, 2011, pp. 1996-1999. 
doi:10.1109/ICASSP.2011.5946902</mixed-citation></ref><ref id="scirp.19576-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">H. Sawada, R. Mukai, S. Araki and S. Makino, “A Robust and Precise Method for Solving the Permutation Problem of Frequency-Domain Blind Source Separation,” IEEE Transactions on Speech and Audio Processing, Vol. 12, No. 5, 2004, pp. 530-538. doi:10.1109/TSA.2004.832994</mixed-citation></ref><ref id="scirp.19576-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">M. S. Pedersen, J. Larsen, U. Kjems and L. C. Parra, “A Survey of Convolutive Blind Source Separation Methods,” Handbook on Speech Processing and Speech Communication, Springer, Berlin, 2007.</mixed-citation></ref><ref id="scirp.19576-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">S. Rickard, “The DUET Blind Source Separation Algorithm,” Blind Speech Separation, Springer, Berlin, 2007.  
doi:10.1007/978-1-4020-6479-1_8</mixed-citation></ref><ref id="scirp.19576-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">A. Jourjine, S. Rickard and O. Yilmaz, “Blind Separation of Disjoint Orthogonal Signals: Demixing n Sources from 2 Mixtures,” IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 5, 2000, pp. 2985-2988. doi:10.1109/ICASSP.2000.861162</mixed-citation></ref><ref id="scirp.19576-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">S. Araki, H. Sawada and S. Makino, “K-Means Based Underdetermined Blind Speech Separation,” Blind Speech Separation, Springer, Berlin, 2007. 
doi:10.1007/978-1-4020-6479-1_9</mixed-citation></ref><ref id="scirp.19576-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">R. O. Duda, P. E. Hart and D. G. Stork, “Pattern Classification,” Wiley &amp; Sons Ltd., New York, 2000.</mixed-citation></ref><ref id="scirp.19576-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">M. S. Pedersen, D. L. Wang, J. Larsen and U. Kjems “Two-Microphone Separation of Speech Mixtures,” IEEE Transactions on Neural Networks, Vol. 19, No. 3, 2008, pp. 475-492. doi:10.1109/TNN.2007.911740</mixed-citation></ref><ref id="scirp.19576-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">D. L. Wang, “On Ideal Binary Mask as the Computational Goal of Auditory Scene Analysis,” Speech Separation by Humans and Machines, Springer, Berlin, 2005, pp. 181-197. doi:10.1007/0-387-22794-6_12</mixed-citation></ref><ref id="scirp.19576-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">I. Jafari, R. Togneri and S. Nordholm, “Review of Multi-Channel Source Separation in Realistic Environments,” 13th Australasian International Conference on Speech Science and Technology, Melbourne, 14-16 December 2010, pp. 201-204.</mixed-citation></ref><ref id="scirp.19576-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">S. Araki and T. Nakatani, “Hybrid Approach for Multichannel Source Separation Combining Time Frequency Mask with Multi-Channel Wiener Filter,” IEEE International Conference on Acoustics, Speech and Signal Processing, Prague, 22-27 May 2011, pp. 225-228.  
doi:10.1109/ICASSP.2011.5946381</mixed-citation></ref><ref id="scirp.19576-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">L.Wang, H. Ding and F. Yin, “Target Speech Extraction in Cocktail Party by Combining Beamforming and Blind Source Separation,” Journal Acoustics Australia, Vol. 39, No. 2, 2011, pp. 64-68.</mixed-citation></ref></ref-list></back></article>