Forecasting S&P 500 Stock Index Using Statistical Learning Models
- 1 Department of Industrial Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA
- 2 Department of Industrial Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA
- 3 Department of Industrial Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA
- 4 Department of Industrial Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA
Abstract
Forecasting the movement of stock market is a long-time attractive topic. This paper implements different statistical learning models to predict the movement of S&P 500 index. The S&P 500 index is influenced by other important financial indexes across the world such as commodity price and financial technical indicators. This paper systematically investigated four supervised learning models, including Logistic Regression, Gaussian Discriminant Analysis (GDA), Naive Bayes and Support Vector Machine (SVM) in the forecast of S&P 500 index. After several experiments of optimization in features and models, especially the SVM kernel selection and feature selection for different models, this paper concludes that a SVM model with a Radial Basis Function (RBF) kernel can achieve an accuracy rate of 62.51% for the future market trend of the S&P 500 index.
- Huang, W., Nakamori, Y. and Wang, S.Y. (2005) Forecasting Stock Market Movement Direction with Support Vector Machine. Computers & Operations Research, 32, 2513-2522. https://doi.org/10.1016/j.cor.2004.03.016
- Choudhry, R. and Garg, K. (2008) A Hybrid Machine Learning System for Stock Market Forecasting. World Academy of Science, Engineering and Technology, 39, 315-318. http://waset.org/publications/8952/a-hybrid-machine-learning-system-for-stock-market-forecasting
- Kim, K. (2003) Financial Time Series Forecasting Using Support Vector Machines. Neurocomputing, 55, 307-319. https://doi.org/10.1016/s0925-2312(03)00372-2
- Quinlan, J.R. (2014) C4. 5: Programs for Machine Learning. Elsevier, 58-60. https://books.google.com/books/about/C4_5.html?id=b3ujBQAAQBAJ
- Bradley, A.P. (1997) The Use of the Area under the ROC Curve in the Evaluation of Machine Learning Algorithms. Pattern Recognition, 30, 1145-1159. https://doi.org/10.1016/S0031-3203(96)00142-2
- Meng, X., Bradley, J., Yuvaz, B., Sparks, E., Venkataraman, S., Liu, D. and Xin, D. (2016) Mllib: Machine Learning in Apache Spark. JMLR, 17, 1-7. http://www.jmlr.org/papers/volume17/15-237/15-237.pdf
- Zaharia, M., Chowdhury, M., Franklin, M.J., Shenker, S. and Stoica, I. (2010) Spark: Cluster Computing with Working Sets. HotCloud, 10, 10-10. http://static.usenix.org/legacy/events/hotcloud10/tech/full_papers/Zaharia.pdf