Low-Rank Sparse Representation with Pre-Learned Dictionaries and Side Information for Singing Voice Separation
- 1 Department of Mathematics, Shanghai University, Shanghai, China
- 2 Department of Mathematics, Shanghai University, Shanghai, China
Abstract
At present, although the human speech separation has achieved fruitful results, it is not ideal for the separation of singing and accompaniment. Based on low-rank and sparse optimization theory, in this paper, we propose a new singing voice separation algorithm called Low-rank, Sparse Representation with pre-learned dictionaries and side Information (LSRi). The algorithm incorporates both the vocal and instrumental spectrograms as sparse matrix and low-rank matrix, meanwhile combines pre-learning dictionary and the reconstructed voice spectrogram form the annotation. Evaluations on the iKala dataset show that the proposed methods are effective and efficient for singing voice separation.
- Li, Y. and Wang, D.L. (2007) Separation of Singing Voice from Music Accompaniment for Monaural Recordings. IEEE Transactions on Audio, Speech and Language Processing, 15, 1475-1487. https://doi.org/10.1109/TASL.2006.889789
- Candes, E.J., Li, X., Ma, Y. and Wright, J. (2011) Robust Principal Component Analysis? Journal of the ACM, 58, 1-37. https://doi.org/10.1145/1970392.1970395
- Huang, P.S., Chen, S.D., Smaragdis, P. and Johnson, M.H. (2012) Singing Voice Separation from Monaural Recordings Using Robust Principal Component Analysis. 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, 25-30 March 2012, 57-60. https://doi.org/10.1109/ICASSP.2012.6287816
- Yu, S., Zhang, H. and Duan, Z. (2017) Singing Voice Separation by Low-Rank and Sparse Spectrogram Decomposition with Pre-Learned Dictionaries. Journal of the Audio Engineering Society, 65, 377-388. https://doi.org/10.17743/jaes.2017.0009
- Chan, T.S., Yeh, T.C., Fan, Z.C., Chen, H.W., Su, L., Yang, Y.H. and Jang, R. (2015) Vocal Activity Informed Singing Voice Separation with the iKala Dataset. 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, 19-24 April 2015, 718-722. https://doi.org/10.1109/ICASSP.2015.7178063
- Lehner, B., Widmer, G. and Sonnleitner, R. (2014) On the Reduction of False Positives in Singing Voice Detection. 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, 4-9 May 2014, 7480-7484. https://doi.org/10.1109/ICASSP.2014.6855054
- Yoshii, K., Fujihara, H., Nakano, T. and Goto, M. (2014) Cultivating Vocal Activity Detection for Music Audio Signals in a Circulation Type Crowd Sourcing Ecosystem. 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, 4-9 May 2014, 624-628. https://doi.org/10.1109/ICASSP.2014.6853671
- Chan, T.S. and Yang, Y.H. (2017) Informed Group-Sparse Representation for Singing Voice Separation. IEEE Signal Processing Letters, 24, 156-160.
- Chan, T.S., Yeh, T.C., Fan, Z.C., Chen, H.W., Sui, L., Yang, Y.H. and Jang, R. (2015) Vocal Activity Informed Singing Voice Separation with the iKala Dataset. 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, 19-24 April 2015, 718-722. https://doi.org/10.1109/ICASSP.2015.7178063
- Boyd, S., Parikh, N., Chu, E., Peleato, B. and Eckstein, J. (2011) Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers. Foundations and Trends in Machine Learning, 3, 1-122. https://doi.org/10.1561/2200000016