Single Channel Source Separation Using Filterbank and 2D Sparse Matrix Factorization
- 1 School of Electrical and Electronic Engineering, Newcastle University, England, UK
- 2 School of Electrical and Electronic Engineering, Newcastle University, England, UK
- 3 School of Electrical and Electronic Engineering, Newcastle University, England, UK
- 4 School of Electrical and Electronic Engineering, Newcastle University, England, UK
- 5 School of Electrical and Electronic Engineering, Newcastle University, England, UK
- 6 Faculty of Information Engineering, Guangdong University of Technology, Guangzhou, China
- 7 School of Marine Science and Technology, Newcastle University, England, UK.
Abstract
We present a novel approach to solve the problem of single channel source separation (SCSS) based on filterbank tech nique and sparse non-negative matrix two dimensional deconvolution (SNMF2D). The proposed approach does not require training information of the sources and therefore, it is highly suited for practicality of SCSS. The major problem of most existing SCSS algorithms lies in their inability to resolve the mixing ambiguity in the single channel observa tion. Our proposed approach tackles this difficult problem by using filterbank which decomposes the mixed signal into sub-band domain. This will result the mixture in sub-band domain to be more separable. By incorporating SNMF2D algorithm, the spectral-temporal structure of the sources can be obtained more accurately. Real time test has been con ducted and it is shown that the proposed method gives high quality source separation performance.
- T. Kristjansson, H. Attias and J. Hershey, “Single Microphone Source Separation Using High Resolution Signal Reconstruction,” Proceedings of International Conference on Acoustics, Speech, and Signal Processing, Quebec, 17-21 May 2004, pp. 817-820.
- B. Gao, W. L. Woo and S. S. Dlay, “Single Channel Source Separation Using EMD-Subband Variable Regularized Sparse Features,” IEEE Transactions on Audio, Speech and Language Processing, Vol. 19, No. 4, 2011, pp. 961-976. doi:10.1109/TASL.2010.2072500
- M. H. Radfa and R. M. Dansereau, “Single-Channel Speech Separation Using Soft Mask Filtering,” IEEE Transactions on Audio, Speech and Language Processing, Vol. 15, No. 8, 2007, pp. 2299-2310.
- D. Ellis, “Model-Based Scene Analysis,” In: D. Wang and G. Brown, Eds. Computational Auditory Scene Analysis: Principles, Algorithms, and Applications, Wiley/ IEEE Press, New York, 2006.
- T. Kristjansson, H. Attias and J. Hershey, “Single MicroPhone Source Separation Using High Resolution Signal Reconstruction,” Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Montreal, 17-21 May 2004, pp. 817-820.
- B. Gao, W. L. Woo and S. S. Dlay, “Adaptive Sparsity NonNegative Matrix Factorization for Single Channel Source Separation,” IEEE Journal of Selected Topics in Signal Processing, Vol. 5, No. 5, 2011, pp. 1932-4553.
- G. J. Brown and M. Cooke, “Computational Auditory Scene Analysis,” Computer Speech and Language, Vol. 8, No. 4, 1994, pp. 297-336. doi:10.1006/csla.1994.1016
- M. Helén and T. Virtanen, “Separation of Drums from Polyphonic Music Using Nonnegative Matrix Factorization and Support Vector Machine,” 13th European Signal Processing Conference, Turkey, 6 September 2005.
- P. Smaragdis and J. C. Brown, “Non-Negative Matrix Factorization for Polyphonic Music Transcription,” IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 19-22 October 2003, pp. 177-180.
- R. Kompass, “A Generalized Divergence Measure for Nonnegative Matrix Factorization,” Proceedings of the Neuroinformatics Workshop, Torun, September 2005.
- A. Cichocki, R. Zdunek, and S. I. Amari, “Csiszár’s divergences for non-negative matrix factorization: family of new algorithms,” Proceedings of the 6th International Conference on Independent Component Analysis and Blind Signal Separation, Vol. 3889, Springer, Charleston, 2006, pp. 32-39. doi:10.1007/11679363_5