Data Integration Techniques for the Construction of an Integrated Database on the Economic Sustainability of Italian Families — Oak Academic Publishing
Research ArticleOpen AccessGoogle Scholar indexed
Data Integration Techniques for the Construction of an Integrated Database on the Economic Sustainability of Italian Families
Department of Humanities Research and Innovation, University of Bari Aldo Moro, Bari, Italy
,
Department of Economics, Management and Business Law, University of Bari Aldo Moro, Bari, Italy
,
Department of Economics, Management and Business Law, University of Bari Aldo Moro, Bari, Italy
,
Department of Humanities Research and Innovation, University of Bari Aldo Moro, Bari, Italy
,
Department of Statistical Sciences, University of Rome La Sapienza, Rome, Italy
1 Department of Humanities Research and Innovation, University of Bari Aldo Moro, Bari, Italy
2 Department of Economics, Management and Business Law, University of Bari Aldo Moro, Bari, Italy
3 Department of Economics, Management and Business Law, University of Bari Aldo Moro, Bari, Italy
4 Department of Humanities Research and Innovation, University of Bari Aldo Moro, Bari, Italy
5 Department of Statistical Sciences, University of Rome La Sapienza, Rome, Italy
This work describes a data integration model using the Statistical Matching methodology (hot deck distance) to integrate two surveys conducted by ISTAT (EU-SILC) and the Bank of Italy (Household Income Survey). The construction of an integrated database based on these surveys is useful for studying consumer behavior in relation to specific commodity groups, analyzing household savings decisions, assessing economic and social inequality, and evaluating the impact of public policies through simulations. The coexistence of multiple and diverse objectives necessitates a highly general and versatile integrated file, providing detailed information on spending patterns, savings levels, income distribution, occupational conditions of household members, and more.
Rässler, S. (2002) Statistical Matching: A Frequentist Theory, Applied Bayesian Methodology and a New Approach. Springer, 266 p.
Christen, P. (2012) Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer, 221 p.
Fellegi, I. (1997) Record Linkage and Public Policy: A Dynamic Evolution. In: Al-Vey, W. and Jamerson, B. Eds., Record Linkage Techniques , Arlington, 3-12.
Belin, T.R. and Rubin, D.B. (1995) A Method for Calibrating False-Match Rates in Record Linkage. Journal of the American Statistical Association , 90, 137-147. https://doi.org/10.2307/2291082
D’Orazio, M., Di Zio, M. and Scanu, M. (2002) Statistical Matching and Official Statistics. Quaderni di Ricerca ISTAT, 1.
Conti, P. and Marella, D. (2014) Uncertainty in Statistical Matching for Complex Sample Surveys. 47 th Meeting of the Italian Statistical Society , Cagliari, 11-13 June 2014, 1-6.
Bohensky, M.A., Jolley, D., Sundararajan, V., Evans, S., Pilcher, D.V., Scott, I., et al. (2010) Data Linkage: A Powerful Research Tool with Potential Problems. BMC Health Services Research , 10, Article No. 346. https://doi.org/10.1186/1472-6963-10-346
UNECE (2021) Guidelines on Statistical Data Integration. United Nations Economic Commission for Europe, 174 p. https://unece.org/info/publications/pub/36715
Schionato, L. (1995) Tecniche di linkage statistico per il raccordo di una pluralità di fonti amministrative. In: Biffignandi, S. and Maritni, M., Eds., Il Registro Statistico Europeo Delle Imprese , Franco Angeli Editore, 333-343.
Winkler, W.E., Yancey, W.E. and Porter, E.H. (2006) The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis. In Advances in Record-Linkage Methodology as Applied to Matching the 1985 Census of Tampa , Florida . U.S. Census Bureau.
Jaro, M.A. (1989) Advances in Record-Linkage Methodology as Applied to Matching the 1985 Census of Tampa, Florida. Journal of the American Statistical Association , 84, 414-420. https://doi.org/10.2307/2289924
Nuccitelli, A., Bosio, F. and Fioriti, L. (2004) L’applicazione reclink per il record linkage: Metodologia implementata e linee guida per la sua utilizzazione. Istat.IT.
D’Orazio, M., Tzavidis, N. and Salvati, N. (2021) Framework for Statistical Matching with Auxiliary Information. Statistical Methods & Applications , 30, 861-885.
Fortini, M., Liseo, B. and Scanu, M. (2002) On Bayesian Record Linkage. Research in Official Statistics , 5, 185-198.
Ruggles, N.N. and Ruggles, R. (1974) A strategy for Merging and Matching Micro Data Sets. Annals of Economic and Social Measurement , 3, 353-371.
Fortunato, E. and Morrone, A. (1999) Approcci micro e macro al linkage dei dati delle indagini ISTAT sulle famiglie. Atti del convegno SIS: Verso i censimenti del 2000, 253-260.
Montrone, A., Perchinunno, P. and de Blasi, R. (2012) Statistical Matching of EU-SILC and HBS Data. Statistica & Applicazioni , 10, 113-130.
Ryu, T. and Eick, C. (1998) A Graph-Based Approach for Discovering Various Kinds of Association Rules. Proceedings of the 1998 ACM Symposium on Applied Computing , Atlanta, 27 February-1 March 1998, 260-267.
Eurostat (2017) Handbook on Data Integration Methods. Eurostat Manuals and Guidelines, 158 p.
Zhang, T., Ramakrishnan, R. and Livny, M. (1996) BIRCH: An Efficient Data Clustering Method for Very Large Databases. ACM SIGMOD Record , 25, 103-114. https://doi.org/10.1145/235968.233324
Everitt, B.S., Landau, S., Leese, M. and Stahl, D. (2011) Cluster Analysis. Wiley. https://doi.org/10.1002/9780470977811
Chiu, T., Fang, D., Chen, J., Wang, Y. and Jeris, C. (2001) A Robust and Scalable Clustering Algorithm for Mixed Type Attributes in Large Database Environment. Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , San Francisco, 26-29 August 2001, 263-268. https://doi.org/10.1145/502512.502549
Huang, Z. (1997) A Fast Clustering Algorithm to Cluster Very Large Categorical Data Sets in Data Mining. Data Mining and Knowledge Discovery , 3, 34-39.
Huang, Z. (1998) Extensions to the K-Means Algorithm for Clustering Large Data Sets with Categorical Values. Data Mining and Knowledge Discovery , 2, 283-304. https://doi.org/10.1023/a:1009769707641
Vathy-Fogarassy, Á. and Abonyi, J. (2013) Graph-Based Clustering and Data Visualization Algorithms. Springer.
Banfield, J.D. and Raftery, A.E. (1993) Model-Based Gaussian and Non-Gaussian Clustering. Biometrics , 49, 803-821. https://doi.org/10.2307/2532201
Vermunt, J.K. and Magidson, J. (2002) Latent Class Cluster Analysis. In: Ha-Genaars, J.A. and Mcphee, I., Eds., Applied Latent Class Analysis , Cambridge University Press, 89-106. https://doi.org/10.1017/cbo9780511499531.004
Schwarz, G. (1978) Estimating the Dimension of a Model. The Annals of Statistics , 6, 461-464. https://doi.org/10.1214/aos/1176344136
Fraley, C. and Raftery, A.E. (2002) Model-Based Clustering, Discriminant Analysis, and Density Estimation. Journal of the American Statistical Association , 97, 611-631. https://doi.org/10.1198/016214502760047131
Pelleg, D. and Moore, A.W. (2000) X-Means: Extending K-Means with Efficient Estimation of the Number of Clusters. Proceedings of the Seventeenth International Conference on Machine Learning ( ICML 2000), Stanford, 29 June 2000-2 July 2000, 727-734.
Steinley, D. (2006) K‐Means Clustering: A Half‐Century Synthesis. British Journal of Mathematical and Statistical Psychology , 59, 1-34. https://doi.org/10.1348/000711005x48266
Whelan, C.T., Layte, R. and Maître, B. (2003) Persistent Income Poverty and Deprivation in the European Union: An Analysis of the First Three Waves of the European Community Household Panel. Journal of Social Policy , 32, 1-18. https://doi.org/10.1017/s0047279402006864
Tanton, R., Vidyattama, Y., Mcnamara, J. and Vu, Q.N. (2009) Small Area Estimation for Local Economic Indicators in Australia. Australasian Journal of Regional Studies , 15, 303-325.
Jenkins, S.P. and van Kerm, P. (2009) The Measurement of Economic Inequality. In: Salverda, W., Nolan, B. and Smeeding, T., Eds., The Oxford Handbook of Economic Inequality , Oxford University Press, 40-67.
Nolan, B. and Whelan, C.T. (2011) Poverty and Deprivation in Europe. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199588435.001.0001
D’Orazio, M., Di Zio, M. and Scanu, M. (2006) Statistical Matching. Wiley. https://doi.org/10.1002/0470023554
Little, R. and Rubin, D. (2019) Statistical Analysis with Missing Data. 3rd Edition, Wiley. https://doi.org/10.1002/9781119482260
OECD (2022) Enhancing the Use of Administrative Data in Official Statistics. OECD Statistics Working Papers, No. 3.
OECD (2019) Measuring the Effectiveness of Social Protection. Social Policy Report, 108.
EUROFOUND (2020) Addressing Household Over-Indebtedness. Publications Office of the European Union, 92 p.
Alkire, et al. (2015) Multidimensional Poverty Measurement and Analysis Get Access Arrow. Oxford University Press.
Atkinson, A.B. (2019) Measuring Poverty around the World. Princeton University Press, 464. https://doi.org/10.2307/j.ctvc77fd6