Completeness Problem of the Deep Neural Networks
- 1 Department of Engineering Technology, Savannah State University, Savannah, Georgia, USA
- 2 Department of Mathematics, Savannah State University, Savannah, Georgia, USA
Abstract
Hornik, Stinchcombe & White have shown that the multilayer feed forward networks with enough hidden layers are universal approximators. Roux & Bengio have proved that adding hidden units yield a strictly improved modeling power, and Restricted Boltzmann Machines (RBM) are universal approximators of discrete distributions. In this paper, we provide yet another proof. The advantage of this new proof is that it will lead to several new learning algorithms. We prove that the Deep Neural Networks implement an expansion and the expansion is complete. First, we briefly review the basic Boltzmann Machine and that the invariant distributions of the Boltzmann Machine generate Markov chains. We then review the θ -transformation and its completeness, i . e. any function can be expanded by θ -transformation. We further review ABM (Attrasoft Boltzmann Machine). The invariant distribution of the ABM is a θ -transformation; therefore, an ABM can simulate any distribution. We discuss how to convert an ABM into a Deep Neural Network. Finally, by establishing the equivalence between an ABM and the Deep Neural Network, we prove that the Deep Neural Network is complete.
- Hinton, G.E., Osindero, S. and Teh, Y. (2006) A Fast Learning Algorithm for Deep Belief Nets. Neural Computation, 18, 1527-1554. https://doi.org/10.1162/neco.2006.18.7.1527
- Bengio, Y. (2009) Learning Deep Architectures for AI (PDF). Foundations and Trends in Machine Learning, 2, 1-127. https://doi.org/10.1561/2200000006
- Amari, S., Kurata, K. and Nagaoka, H. (1992) Information Geometry of Boltzmann Machine. IEEE Transactions on Neural Networks, 3, 260-271. https://doi.org/10.1109/72.125867
- Byrne, W. (1992) Alternating Minimization and Boltzmann Machine Learning. IEEE Transactions on Neural Networks, 3, 612-620. https://doi.org/10.1109/72.143375
- Coursera (2017). https://www.coursera.org/
- TensorFlow (2017). https://www.tensorflow.org/
- Torch (2017). http://torch.ch/
- Theano (2017). http://deeplearning.net/software/theano/introduction.html
- Hornik, K., Stinchcombe, M. and White, H. (1989) Multilayer Feedforward Networks Are Universal Approximators. Neural Networks, 2, 359-366. https://doi.org/10.1016/0893-6080(89)90020-8
- Le Roux, N. and Bengio, Y. (2008) Representational Power of Restricted Boltzmann Machines and Deep Belief Networks. Neural Computation, 20, 1631-1649. https://doi.org/10.1162/neco.2008.04-07-510
- Feller, W. (1968) An Introduction to Probability Theory and Its Application. John Wiley and Sons, New York.
- Liu, Y. (1993) Image Compression Using Boltzmann Machines. Proceedings of SPIE, 2032, 103-117. https://doi.org/10.1117/12.162027
- Liu, Y. (1995) Boltzmann Machine for Image Block Coding. Proceedings of SPIE, 2424, 434-447. https://doi.org/10.1117/12.205245
- Liu, Y. (1997) Character and Image Recognition and Image Retrieval Using the Boltzmann Machine. Proceedings of SPIE, 3077, 706-715. https://doi.org/10.1117/12.271533
- Liu, Y. (2002) US Patent No. 7, 773, 800. http://www.google.com/patents/US7773800