Business Data Extraction Using a Programming Language
- 1 School of Economics, Business Administration and Accounting at Ribeirão Preto, University of São Paulo, São Paulo, Brazil
- 2 School of Economics, Business Administration and Accounting at Ribeirão Preto, University of São Paulo, São Paulo, Brazil
- 3 School of Economics, Business Administration and Accounting at Ribeirão Preto, University of São Paulo, São Paulo, Brazil
- 4 School of Economics, Business Administration and Accounting at Ribeirão Preto, University of São Paulo, São Paulo, Brazil
Abstract
In the era of great informational quantity, the presence of technologies that assist in the extraction, transformation, and loading of data has become increasingly necessary. The term Big Data, usually used to describe this volume of information, requires the user to have knowledge of multiple tools such as Excel, VBA, SQL, Tableau, Python, Spark, AWS, and so on. In this context, t he present work aims to study data extraction techniques using different methodologies. At the end of the work, a library of functions in the Python language will be made available that will deliver a compilation of stock price information available on the Yahoo Finance website as well as balance sheets from financial institutions released by Bacen. The main resource used will be Web Scraping, which is a method that aims to automate data collection via the web. Once the collection of functions has been structured, it will be made available for public enjoyment through the GitHub platform.
- Anaconda (2020) Individual Edition. https://www.anaconda.com/products/individual
- Baumgartner, R., Frohlich, O., Gottlob, G., Harz, P., Herzog, M., & Lehmann, P. (2005). Web Data Extraction for Business Intelligence: The Lixto Approach. Gesellschaft für Informatik eV.
- Borges, L. E. (2014). Python para desenvolvedores: Aborda Python 3.3 (p. 14). Novatec Editora. https://books.google.com.br/books?hl=pt-BR&lr=lang_pt&id=eZmtBAAAQBAJ&oi=fnd&pg= PA14&dq=python&ots=VDSosqIkmu&sig=TZ0MbKn058lnRJ9zgZrNLmoOFh4#v=onepage&q=python&f=false
- Bradley, E. H., Curry, L. A., & Devers, K. J. (2007). Qualitative Data Analysis for Health Services Research: Developing Taxonomy, Themes, and Theory. Health Services Research, 42, 1758-1772. https://doi.org/10.1111/j.1475-6773.2006.00684.x
- Buriol, T. M., Marco, B., & Argenta, M. A. (2009). Acelerando o desenvolvimento eo processamento de análises numéricas computacionais utilizando python e cuda. https://www.researchgate.net/profile/Marco_Argenta/publication/228683446_ACELERANDO_ O_DESENVOLVIMENTO_EO_PROCESSAMENTO_DE_ANALISES_NUMERICAS_COMPUTACIONAIS_UTILIZANDO_ PYTHON_E_CUDA/links/5630d6c908ae0530378cdf06.pdf
- Catanese, S. A., De Meo, P., Ferrara, E., Fiumara, G., & Provetti, A. (2011, May). Crawling Facebook for Social Network Analysis Purposes. In Proceedings of the International Conference on Web Intelligence, Mining and Semantics (pp. 1-8). https://doi.org/10.1145/1988688.1988749
- Chen, H., Chau, M., & Zeng, D. (2002). CI Spider: A Tool for Competitive Intelligence on the Web. Decision Support Systems, 34, 1-17. https://doi.org/10.1016/S0167-9236(02)00002-7
- Chen, M., Mao, S., & Liu, Y. (2014). Big Data: A Survey. Mobile Networks and Applications, 19, 171-209. https://doi.org/10.1007/s11036-013-0489-0
- Ferrara, E., De Meo, P., Fiumara, G., & Baumgartner, R. (2014). Web Data Extraction, Applications and Techniques: A Survey. Knowledge-Based Systems, 70, 301-323. https://doi.org/10.1016/j.knosys.2014.07.007
- Github (2020) Fabiobragato/Finances. https://github.com/fabiobragato/finances.git
- Gkotsis, G., Stepanyan, K., Cristea, A. I., & Joy, M. (2013, July). Self-Supervised Automated Wrapper Generation for Weblog Data Extraction. In British National Conference on Databases (pp. 292-302). Springer. https://doi.org/10.1007/978-3-642-39467-6_26
- Laender, A. H., Ribeiro-Neto, B. A., Da Silva, A. S., & Teixeira, J. S. (2002). A Brief Survey of Web Data Extraction Tools. ACM SIGMOD Record, 31, 84-93. https://doi.org/10.1145/565117.565137