Uncovering and Displaying the Coherent Groups of Rank Data by Exploratory Riffle Shuffling
- 1 Université de Moncton, Moncton, Canada
- 2 Université de Moncton, Moncton, Canada
Abstract
Let n respondents rank order d items, and suppose that . Our main task is to uncover and display the structure of the observed rank data by an exploratory riffle shuffling procedure which sequentially decomposes the n voters into a finite number of coherent groups plus a noisy group: where the noisy group represents the outlier voters and each coherent group is composed of a finite number of coherent clusters. We consider exploratory riffle shuffling of a set of items to be equivalent to optimal two blocks seriation of the items with crossing of some scores between the two blocks. A riffle shuffled coherent cluster of voters within its coherent group is essentially characterized by the following facts: 1) Voters have identical first TCA factor score, where TCA designates taxicab correspondence analysis, an L 1 variant of corresponden ce analysis; 2) Any preference is easily interpreted as riffle shuffling of its items; 3) The nature of different riffle shuffling of items can be seen in the structure of the contingency table of the first-order marginals constructed from the Borda scorings of the voters; 4) The first TCA factor scores of the items of a coherent cluster are interpreted as Borda scale of the items. We also introduce a crossing index, which measures the extent of crossing of scores of voters between the two blocks seriation of the items. The novel approach is explained on the benchmarking SUSHI data set, where we show that this data set has a very si mple structure, which can also be communicated in a tabular form.
- Kamishima, T. (2003) Nantonac Collaborative Filtering: Recommendation Based on Order Responses. Proceedings of the 9th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Washington DC, August 2003, 583-588. https://doi.org/10.1145/956750.956823
- Huang, J. and Guestrin, C. (2012) Uncovering the Riffled Independence Structure of Ranked Data. Electronic Journal of Statistics, 6, 199-230. https://doi.org/10.1214/12-EJS670
- Lu, T. and Boutilier, C. (2014) Effective Sampling and Learning for Mallows Models with Pairwise Preference Data. Journal of Machine Learning Research, 15, 3783-3829.
- Vitelli, V., Sørenson, Ø., Crispino, M., Frigessi, A. and Arjas, E. (2018) Probabilistic Preference Learning with the Mallows Rank Model. Journal of Machine Learning Research, 18, 1-49.
- Diaconis, P. (1989) A Generalization of Spectral Analysis with Application to Ranked Data. Annals of Statistics, 17, 949-979. https://doi.org/10.1214/aos/1176347251
- Marden, J.I. (1995) Analyzing and Modeling of Rank Data. Chapman & Hall, London
- Alvo, M. and Yu, P. (2014) Statistical Methods for Ranking Data. Springer, New York. https://doi.org/10.1007/978-1-4939-1471-5
- Bayer, D. and Diaconis, P. (1992) Trailing the Dovetail Shuffle to Its Lair. Annals of Probability, 2, 294-313. https://doi.org/10.1214/aoap/1177005705
- Choulakian, V. (2016) Globally Homogenous Mixture Components and Local Heterogeneity of Rank Data. arXiv:1608.05058
- Choulakian, V. (2006) Taxicab Correspondence Analysis. Psychometrika, 71, 333-345. https://doi.org/10.1007/s11336-004-1231-4
- Choulakian, V. (2016) Matrix Factorizations Based on Induced Norms. Statistics, Optimization and Information Computing, 4, 1-14. https://doi.org/10.19139/soic.v4i1.160
- De Borda, J. (1781) Mémoire sur les élections au scrutin. Histoire de L’Académie Royale des Sciences, 102, 657-665.
- Benzécri, J.P. (1991) Comment on Leo A. Goodman’s Invited Paper. Journal of the American Statistical Association, 86, 1112-1115. https://doi.org/10.1080/01621459.1991.10475157
- Van de Velden, M. (2000) Dual Scaling and Correspondence Analysis of Rank Order Data. In: Heijmans, R.D.H., Pollock, D.S.G. and Satorra, A. Eds., Innovations in Multivariate Statistical Analysis, Vol. 36, Kluwer Academic Publishers, Dordrecht, 87-99. https://doi.org/10.1007/978-1-4615-4603-0_6