Sound-Environment Monitoring Method Based on Computational Auditory Scene Analysis
- 1 Service Sensing, Assimilation, and Modeling Research Group, Human Informatics Research Institute, National Institute of Advanced Industrial Science and Technology (AIST), Tsukuba, Japan
Abstract
Monitoring techniques are a key technology for examining the conditions in various scenarios, e.g., structural conditions, weather conditions, and disasters. In order to understand such scenarios, the appropriate extraction of their features from observation data is important. This paper proposes a monitoring method that allows sound environments to be expressed as a sound pattern. To this end, the concept of synesthesia is exploited. That is, the keys, tones, and pitches of the monitored sound are expressed using the three elements of color, that is, the hue, saturation, and brightness, respectively. In this paper, it is assumed that the hue, saturation, and brightness can be detected from the chromagram, sonogram, and sound spectrogram, respectively, based on a previous synesthesia experiment. Then, the sound pattern can be drawn using color, yielding a “painted sound map.” The usefulness of the proposed monitoring technique is verified using environmental sound data observed at a galleria.
- Hamamoto, T. (2015) Structural Health Monitoring of Buildings. Transactions of Foundation Engineering & Equipment, 43, 17-20. (In Japanese)
- Chachada, J.S. and Kuo, C.-C.J. (2014) Environmental Sound Recognition: A Survey. SIP (2014), Vol. 3, e14, 1-15. https://www.cambridge.org/core/services/aop-cambridge-core/content/view/S2048770314000122
- Mitrovic, D., Zeppelzauer, M. and Breiteneder, C. (2010) Features for Content-Based Audio Retrieval. In: Advances in Computers, Vol. 78, Elsevier, Amsterdam, 71-150.
- Deng, J.D., Simmermacher, C. and Cranefield, S. (2008) A Study on Feature Analysis for Musical Instrument Classification. IEEE Transactions on Systems, Man, and Cybernetics, Part B, 38, 429-438. https://doi.org/10.1109/TSMCB.2007.913394
- Peltonen, V., Tuomi, J., Klapuri, A., Huopaniemi, J. and Sorsa, T. (2002) Computational Auditory Scene Recognition. 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Orlando, FL, 13-17 May 2002, II-1941-II-1944.
- Potamitis, I. and Ganchev, T. (2008) Generalized Recognition of Sound Events: Approaches and Applications. In: Tsihrintzis, G.A. and Jain, L.C., Eds., Multimedia Services in Intelligent Environments, Springer, Berlin, Heidelberg, 41-79.
- Wang, J.-C., Wang, J.-F., He, K.W. and Hsu, C.-S. (2006) Environmental Sound Classification Using Hybrid SVM/KNN Classifier and MPEG-7 Audio Low-Level Descriptor. International Joint Conference on Neural Networks, Vancouver, 16-21 July 2006, 1731-1735.
- Muhammad, G., Alotaibi, Y.A., Alsulaiman, M. and Huda, M.N. (2010) Environment Recognition Using Selected MPEG-7 Audio Features and Mel-Frequency Cepstral Coefficients. 2010 5th International Conference on Digital Telecommunications (ICDT), Athens, 13-19 June 2010, 11-16. https://doi.org/10.1109/ICDT.2010.10
- Tsau, E., Kim, S.-H. and Kuo, C.-C.J. (2011) Environmental Sound Recognition with CELP-Based Features. 2011 10th International Symposium on Signals, Circuits and Systems (ISSCS), lasi, 30 June-1 July 2011, 1-4. https://doi.org/10.1109/ISSCS.2011.5978729
- Chu, S., Narayanan, S. and Kuo, C.-C.J. (2009) Environmental Sound Recognition with Time-Frequency Audio Features. IEEE Transactions on Audio, Speech, and Language Processing, 17, 1142-1158. https://doi.org/10.1109/TASL.2009.2017438
- The Color Science Association of Japan (2011) Handbook of Color Science. 3rd Edition, University of Tokyo Press, Japan. (In Japanese)