Maximum likelihood estimation for ordinal randomized response data via the EM algorithm: identifiability and efficiency

Authors

  • A. F. Akintola Department of Mathematics and Statistics, Redeemer’s University, Ede, Osun State, Nigeria
  • O. M. Olayiwola Department of Statistics, College of Physical Sciences, Federal University of Agriculture, Abeokuta, Ogun State, Nigeria
  • A. A. Akintunde Department of Statistics, College of Physical Sciences, Federal University of Agriculture, Abeokuta, Ogun State, Nigeria
  • I. A. Osinuga Department of Mathematics, College of Physical Sciences, Federal University of Agriculture, Abeokuta, Ogun State, Nigeria

Keywords:

Randomized response, Ordinal data, Identifiability, EM algorithm, Maximum likelihood

Abstract

Surveys on sensitive matters, including security behaviour, substance use, and financial misconduct, are prone to insincere answers when respondents fear disapproval or consequences. Randomized response techniques address this by letting respondents answer through a privacy-preserving device, so no individual reply is ever revealed. The approach is well developed for dichotomous and quantitative questions but has seldom been adapted to ordinal Likert-type items, among the most common measurement tools in behavioural research. This article provides that extension. We propose a scrambling mechanism for four-category ordinal responses driven by a single six-sided die and show that the design is identifiable: the population category probabilities are uniquely determined by the population distribution of scrambled responses, and this can be checked before any data are collected. Because the true category is latent, we derive an expectation-maximization (EM) algorithm with closed-form updates that converges to the unique maximiser of the observed-data likelihood, and we give the asymptotic theory in the three free coordinates the simplex actually has. The maximum-likelihood and method-of-moments estimators are asymptotically equivalent, but they differ appreciably in finite samples. Across five distributions and sample sizes up to 1,000, the maximum-likelihood estimator reduces mean squared error by up to 27% at (n=200) for polarized distributions and by construction avoids the boundary violations that force the moment estimator to be projected back onto the simplex in over half of small polarized samples. The result is a practical and reliable tool for measuring sensitive ordinal outcomes without compromising respondent trust.

0 0

References

[1] R. Tourangeau & T. Yan, “Sensitive questions in surveys”, Psychological bulletin 133 (2007) 859. https://doi.org/10.1037/0033-2909.133.5.859.

[2] S. L. Warner, “Randomized response: a survey technique for eliminating evasive answer bias”, Journal of the American statistical association 60 (1965) 63. https://doi.org/10.1080/01621459.1965.10480775.

[3] A. Chaudhuri & R. Mukerjee, Randomized response: theory and techniques, Marcel Dekker, New York, USA, 1988. https://archive.org/details/randomizedrespon0000chau.

[4] G. J. L. M. Lensvelt-Mulders, J. J. Hox, P. G. M. van der Heijden & C. J. M. Maas, “Meta-analysis of randomized response research: thirty-five years of validation”, Sociological methods & research 33 (2005) 319. https://doi.org/10.1177/0049124104268664.

[5] R. F. Boruch, “Assuring confidentiality of responses in social research: a note on strategies”, The American sociologist 6 (1971) 308. https://www.jstor.org/stable/27701807.

[6] J. A. Fox & P. E. Tracy, Randomized response: a method for sensitive surveys, SAGE Publications, Beverly Hills, USA, 1986.

[7] B. H. Eichhorn & L. S. Hayre, “Scrambled randomized response methods for obtaining sensitive quantitative data”, Journal of statistical planning and inference 7 (1983) 307. https://doi.org/10.1016/0378-3758(83)90002-2.

[8] Diana, G. & Perri, P. F., “ New scrambled response models for estimating the mean of a sensitive quantitative character”, Journal of applied statistics, 37 (2010) 11, 1875-1890. https://doi.org/10.1080/02664760903186031.

[9] B. G. Greenberg, A. L. A. Abul-Ela, W. R. Simmons & D. G. Horvitz, “The unrelated question randomized response model: theoretical framework”, Journal of the American statistical association 64 (1969) 520. https://doi.org/10.1080/01621459.1969.10500991.

[10] Moors, J. J. A. , “Optimization of the unrelated question randomized response model”, Journal of the American statistical association 66 (1971) 629. https://doi.org/10.1080/01621459.1971.10482320.

[11] A. P. Dempster, N. M. Laird & D. B. Rubin, “Maximum likelihood from incomplete data via the EM algorithm”, Journal of the Royal statistical society series B 39 (1977) 1. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x.

[12] C. F. J. Wu, “On the convergence properties of the EM algorithm”, The Annals of statistics 11 (1983) 95. https://doi.org/10.1214/aos/1176346060.

[13] W. Wang & M. Á. Carreira-Perpiñán, “Projection onto the probability simplex: an efficient algorithm with a simple proof, and an application”, arXiv (2013). https://arxiv.org/abs/1309.1541.

[14] M. A. Yunusa, A. Audu, U. Usman, K. O. Aremu & M. Aphane, “Calibrated-two optional randomized response techniques (C-TORRT) for the estimation of quantitative sensitive variable information”, PLOS ONE 21 (2026) e0339271. https://doi.org/10.1371/journal.pone.0339271.

[15] M. Rueda, B. Cobo & P. F. Perri, “Randomized response estimation in multiple frame surveys”, International journal of computer mathematics 97 (2020) 189. https://doi.org/10.1080/00207160.2018.1476856.

[16] J. W. Yu, G. L. Tian & M. L. Tang, “Two new models for survey sampling with sensitive characteristic: design and analysis”, Metrika 67 (2008) 251. https://doi.org/10.1007/s00184-007-0131-x.

[17] M. L. Tang, Q. Wu, G. L. Tian & J. H. Guo, “Two-sample non randomized response techniques for sensitive questions”, Communications in statistics, theory and methods 43 (2014) 408. https://doi.org/10.1080/03610926.2012.657323.

[18] Azeem, M., Asadullah, Ijaz, M., Hussain, S., Salahuddin, N., & Salam, A., “A novel randomized scrambling technique for mean estimation of a finite population”, Heliyon 10 (2024) 11. https://doi.org/10.1016/j.heliyon.2024.e31690.

FIG3

Published

2026-10-02

How to Cite

Maximum likelihood estimation for ordinal randomized response data via the EM algorithm: identifiability and efficiency. (2026). African Scientific Reports, 5(3), 602. https://doi.org/10.46481/asr.2026.5.3.602

Issue

Section

MATHEMATICS AND STATISTICS SECTION

How to Cite

Maximum likelihood estimation for ordinal randomized response data via the EM algorithm: identifiability and efficiency. (2026). African Scientific Reports, 5(3), 602. https://doi.org/10.46481/asr.2026.5.3.602

Similar Articles

11-20 of 132

You may also start an advanced similarity search for this article.