Square Neurons, Power Neurons, and Their Learning Algorithms
- 1 Department of Engineering Technology, Savannah State University, Savannah, Georgia
Abstract
In this paper, we introduce the concepts of square neurons, power neu-rons, and new learning algorithms based on square neurons, and power neurons. First, we briefly review the basic idea of the Boltzmann Machine, specifically that the invariant distributions of the Boltzmann Machine generate Markov chains. We further review ABM (Attrasoft Boltzmann Machine). Next, we review the <i>θ</i>-transformation and its completeness, <i> i.e. </i> any function can be expanded by <i> θ </i> -transformation. The invariant distribution of the ABM is a <i> θ </i> -transformation; therefore, an ABM can simulate any distribution. We review the linear neurons and the associated learning algorithm. We then discuss the problems of the exponential neurons used in ABM, which are unstable, and the problems of the linear neurons, which do not discriminate the wrong answers from the right answers as sharply as the exponential neurons. Finally, we introduce the concept of square neurons and power neurons. We also discuss the advantages of the learning algorithms based on square neurons and power neurons, which have the stability of the linear neurons and the sharp discrimination of the exponential neurons.
- Hinton, G.E., Osindero, S. and Teh, Y. (2006) A Fast Learning Algorithm for Deep Belief Nets. Neural Computation, 18, 1527-1554. https://doi.org/10.1162/neco.2006.18.7.1527
- TensorFlow. https://www.tensorflow.org/
- Torch. http://torch.ch/
- Theano. http://deeplearning.net/software/theano/introduction.html
- Byrne, W. (1992) Alternating Minimization and Boltzmann Machine Learning. IEEE Transactions on Neural Networks, 3, 612-620. https://doi.org/10.1109/72.143375
- Jolliffe, I.T. (2002) Principal Component Analysis, Series: Springer Series in Statistics. 2nd Edition, Springer, New York, 487 p.
- Abdi, H. and Williams, L.J. (2010) Principal Component Analysis. Wiley Interdisciplinary Reviews: Computational Statistics, 2, 433-459.
- Olshausen, B.A. (1996) Emergence of Simple-Cell Receptive Field Properties by Learning a Sparse Code for Natural Images. Nature, 381, 607-609. https://doi.org/10.1038/381607a0
- Gupta, N. and Stopfer, M. (2014) A Temporal Channel for Information in Sparse Sensory Coding. Current Biology, 24, 2247-2256. https://doi.org/10.1016/j.cub.2014.08.021
- Bengio, Y. (2009) Learning Deep Architectures for AI (PDF). Foundations and Trends in Machine Learning, 2, No. 1. https://doi.org/10.1561/2200000006
- Liou, C.-Y., Huang, J.-C. and Yang, W.-C. (2008) Modeling Word Perception Using the Elman Network. Neurocomputing, 71, 3150-3157. https://doi.org/10.1016/j.neucom.2008.04.030
- Liou, C.-Y., Cheng, C.-W., Liou, J.-W. and Liou, D.-R. (2014) Autoencoder for Words. Neurocomputing, 139, 84-96.
- Rumelhart, D.E., Hinton, G.E. and Williams, R.J. (1986) Learning Internal Representations by Error Propagation.
- Bourlard, H. and Kamp, Y. (1988) Auto-Association by Multilayer Perceptrons and Singular Value Decomposition. Biological Cybernetics, 59, 291-294. https://doi.org/10.1007/BF00332918
- Hinton and Salakhutdinov. (2006) Reducing the Dimensionality of Data with Neural Networks.
- MacQueen, J.B. (1967) Some Methods for Classification and Analysis of Multivariate Observations. Proceedings of 5th Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, Vol. 1, 281-297.
- Steinhaus, H. (1957) Sur la division des corps matériels en parties. Bulletin L’Académie Polonaise des Science, 4, 801-804. (In French)