Due to the presence of non-stationarities and discontinuities in the audio signal, segmentation and classification of audio signal is a really challenging task. Automatic music classification and annotation is still considered as a challenging task due to the difficulty of extracting and selecting the optimal audio features. Hence, this paper proposes an efficient approach for segmentation, feature extraction and classification of audio signals. Enhanced Mel Frequency Cepstral Coefficient (EMFCC)-Enhanced Power Normalized Cepstral Coefficients (EPNCC) based feature extraction is applied for the extraction of features from the audio signal. Then, multi-level classification is done to classify the audio signal as a musical or non-musical signal. The proposed approach achieves better performance in terms of precision, Normalized Mutual Information (NMI), F-score and entropy. The PNN classifier shows high False Rejection Rate (FRR), False Acceptance Rate (FAR), Genuine Acceptance rate (GAR), sensitivity, specificity and accuracy with respect to the number of classes.
KeywordsAudio SignalEnhanced Mel Frequency Cepstral Coefficient (EMFCC)Enhanced Power Normalized Cepstral Coefficients (EPNCC)Probabilistic Neural Network (PNN) Classifier
Castán, D., Tavarez, D., Lopez-Otero, P., Franco-Pedroso, J., Delgado, H., Navas, E., et al. (2015) Albayzín-2014 evaluation: Audio segmentation and classification in broadcast news domains. EURASIP Journal on Audio, Speech, and Music Processing, 2015, 33. http://dx.doi.org/10.1186/s13636-015-0076-3
Kubala, F., Jin, H., Matsoukas, S., Nguyen, L., Schwartz, R. and Makhoul, J. (1997) The 1996 BBN Byblos HUB-4 Transcription System. Proceedings of the 1997 DARPA Speech Recognition Workshop, Chantilly, VA, 2-5 February 1997, 90-93.
Bakis, R., Chen, S., Gopalakrishnan, P., Gopinath, R., Maes, S., Polymenakos, L. and Franz, M. (1997) Transcription of Broadcast News Shows with the IBM Large Vocabulary Speech Recognition System. Proceedings of the Speech Recognition Workshop, Chantilly, February 1997, 67-72.
Beigi, H.S. and Maes, S. (1998) Speaker, Channel and Environment Change Detection. Proceedings of the World Congress on Automation, Anchorage, AK, 18 May 1998, 18-22.
Siegler, M.A., Jain, U., Raj, B. and Stern, R.M. (1997) Automatic Segmentation, Classification and Clustering of Broadcast News Audio. Proceedings of DARPA Speech Recognition Workshop, Chantilly, VA, 2-5 February 1997, 97- 99.
Zhang, X., Su, Z., Lin, P., He, Q. and Yang, J. (2014) An Audio Feature Extraction Scheme Based on Spectral Decomposition. International Conference on Audio, Language and Image Processing (ICALIP), Shanghai, 7-9 July 2014, 730-733.
Patil, H.A., Madhavi, M.C., Jain, R. and Jain, A.K. (2012) Combining Evidence from Temporal and Spectral Features for Person Recognition Using Humming. In: Kundu, M.K., Mitra, S., Mazumdar, D. and Pal, S.K., Eds., Perception and Machine Intelligence, Springer, Berlin Heidelberg, 321-328. http://dx.doi.org/10.1007/978-3-642-27387-2_40
Bhalke, D., Rao, C. and Bormane, D.S. (2014) Musical Instrument Classification Using Higher Order Spectra. International Conference on Signal Processing and Integrated Networks (SPIN), Noida, 20-21 February 2014, 40-45. http://dx.doi.org/10.1109/spin.2014.6776918
Haque, M.A. and Kim, J.-M. (2013) An Enhanced Fuzzy C-Means Algorithm for Audio Segmentation and Classification. Multimedia Tools and Applications, 63, 485-500. http://dx.doi.org/10.1007/s11042-011-0921-z
Lude?a-Choez, J. and Gallardo-Antolín, A. (2015) Feature Extraction Based on the High-Pass Filtering of Audio Signals for Acoustic Event Classification. Computer Speech & Language, 30, 32-42. http://dx.doi.org/10.1016/j.csl.2014.04.001
Lefèvre, S. and Vincent, N. (2011) A Two Level Strategy for Audio Segmentation. Digital Signal Processing, 21, 270- 277. http://dx.doi.org/10.1016/j.dsp.2010.07.003
Evangelista, T.L., Priolli, T.M., Silla, C.N., Angelico, B. and Kaestner, C. (2014) Automatic Segmentation of Audio Signals for Bird Species Identification. IEEE International Symposium on Multimedia (ISM), Taichung, 10-12 December 2014, 223-228. http://dx.doi.org/10.1109/ism.2014.46
Dhanalakshmi, P., Palanivel, S. and Ramalingam, V. (2011) Classification of Audio Signals Using AANN and GMM. Applied Soft Computing, 11, 716-723. http://dx.doi.org/10.1016/j.asoc.2009.12.033
Haque, M.A. and Kim, J.-M. (2013) An Analysis of Content-Based Classification of Audio Signals Using a Fuzzy C-Means Algorithm. Multimedia Tools and Applications, 63, 77-92. http://dx.doi.org/10.1007/s11042-012-1019-y
Dhanalakshmi, P., Palanivel, S. and Ramalingam, V. (2011) Pattern Classification Models for Classifying and Indexing Audio Signals. Engineering Applications of Artificial Intelligence, 24, 350-357. http://dx.doi.org/10.1016/j.engappai.2010.10.011
Gergen, S., Nagathil, A. and Martin, R. (2015) Classification of Reverberant Audio Signals Using Clustered Ad Hoc Distributed Microphones. Signal Processing, 107, 21-32. http://dx.doi.org/10.1016/j.sigpro.2014.04.034
Bhat, A.S., Amith, V., Prasad, N.S. and Mohan, D.M. (2014) An Efficient Classification Algorithm for Music Mood Detection in Western and Hindi Music Using Audio Feature Extraction. 5th International Conference on Signal and Image Processing (ICSIP), Jeju Island, 8-10 January 2014, 359-364. http://dx.doi.org/10.1109/icsip.2014.63
Gergen, S. and Martin, R. (2014) Linear Combining of Audio Features for Signal Classification in Ad-Hoc Microphone Arrays. 11 ITG Symposium; Proceedings of Speech Communication, Erlangen, 24-26 September 2014, 1-4.
Pesek, M., Leonardis, A. and Marolt, M. (2014) Boosting Audio Chord Estimation Using Multiple Classifiers. International Conference on Systems, Signals and Image Processing (IWSSIP), Dubrovnik, 12-15 May 2014, 107-110.
Srinivasa Murthy, Y. and Koolagudi, S.G. (2015) Classification of Vocal and Non-Vocal Regions from Audio Songs Using Spectral Features and Pitch Variations. IEEE 28th Canadian Conference on Electrical and Computer Engineering (CCECE), Halifax, 3-6 May 2015, 1271-1276. http://dx.doi.org/10.1109/ccece.2015.7129461
Koolagudi, S.G. and Krothapalli, S.R. (2012) Emotion Recognition from Speech Using Sub-Syllabic and Pitch Synchronous Spectral Features. International Journal of Speech Technology, 15, 495-511. http://dx.doi.org/10.1007/s10772-012-9150-8
Geiger, J.T., Schuller, B. and Rigoll, G. (2013) Large-Scale Audio Feature Extraction and SVM for Acoustic Scene Classification. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, 20-23 October 2013, 1-4. http://dx.doi.org/10.1109/waspaa.2013.6701857
Oh, S.Y. and Chung, K.-Y. (2014) Target speech Feature Extraction Using Non-Parametric Correlation Coefficient. Cluster Computing, 17, 893-899. http://dx.doi.org/10.1007/s10586-013-0284-5
Gaj?ek, R., Miheli?, F. and Dobri?ek, S. (2013) Speaker State Recognition Using an HMM-Based Feature Extraction Method. Computer Speech & Language, 27, 135-150. http://dx.doi.org/10.1016/j.csl.2012.01.007
Anguera, X. (2012) Speaker Independent Discriminant Feature Extraction for Acoustic Pattern-Matching. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, 25-30 March 2012, 485-488. http://dx.doi.org/10.1109/icassp.2012.6287922
Salamon, J., Rocha, B. and Gómez, E. (2012) Musical Genre Classification Using Melody Features Extracted from Polyphonic Music Signals. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, 25-30 March 2012, 81-84. http://dx.doi.org/10.1109/icassp.2012.6287822
Alam, M.J., Kinnunen, T., Kenny, P., Ouellet, P. and O’Shaughnessy, D. (2013) Multitaper MFCC and PLP Features for Speaker Verification Using i-Vectors. Speech Communication, 55, 237-251. http://dx.doi.org/10.1016/j.specom.2012.08.007
Muthumari, A. and Mala, K. (2015) Computerized Methods for Audio Segmentation and Classification: Survey. International Journal of Applied Engineering Research, 10, 26857-26870.
Rajeswari, K.C. and Uma Maheswari, P. (2015) Feature Extraction and Analysis of Speech Quality for Tamil Text System using Fast Fourier Transform. Australian Journal of Basic and Applied Sciences, 9, 349-356.
(2015) 21 August 2015. http://www.dtic.upf.edu/~ffuhrmann/PhD/data/
(2015) 21 August 2015. http://marsyasweb.appspot.com/download/data_sets/
Ji, X., Han, J., Jiang, X., Hu, X., Guo, L., Han, J., et al. (2015) Analysis of Music/Speech via Integration of Audio Content and Functional Brain Response. Information Sciences, 297, 271-282. http://dx.doi.org/10.1016/j.ins.2014.11.020
Su, L., Yeh, C.-C.M., Liu, J.-Y., Wang, J.-C. and Yang, Y.-H. (2014) A Systematic Evaluation of the Bag-of-Frames Representation for Music Information Retrieval. IEEE Transactions on Multimedia, 16, 1188-1200. http://dx.doi.org/10.1109/TMM.2014.2311016
Fu, Z., Lu, G., Ting, K.M. and Zhang, D. (2011) Music Classification via the Bag-of-Features Approach. Pattern Recognition Letters, 32, 1768-1777. http://dx.doi.org/10.1016/j.patrec.2011.06.026
Fuhrmann, F. (2012) Automatic Musical Instrument Recognition from Polyphonic Music Audio Signals. PhD Thesis, Universitat Pompeu Fabra, Barcelona.