Project Activities
The project team addressed this by first developing the theoretical basis for a software package by combining Bayesian, classical, and information-theoretic perspectives into a unified approach to statistical modeling. Then, after extensive Monte Carlo simulations to generate and test the unified approach, they developed a user-friendly R package for CoSME (comprehensive statistical model evaluation).
Key outcomes
The main findings of this project are as follows:
- Goodness-of-fit testing is the primary method of model evaluation in statistical modeling in education research, but provides limited scientific value (Bonifay, 2022; Bonifay & Depaoli, 2023; Bonifay et al., 2024; Bonifay et al., 2025; Cai, 2025).
- Alternative methods like Bayesian predictive model checking are much more insightful, especially in the context of evaluating item response theory models and replicating factor model-based research (Bonifay & Depaoli, 2023; Bonifay et al., 2024; Winter et al., 2025).
- Some models have an inherent tendency to fit better than other models to any possible data set, including, for example, the highly popular p-factor model of general psychopathology (Bonifay, 2022; Bonifay et al., 2025; Suh et al., 2025; Watts et al., 2023).
- The information-theoretic principle of minimum description length is especially useful for evaluating statistical models, as it quantifies the evidential value of good fit (Bonifay, 2022; Bonifay et al, 2025; Suh et al., 2025).
People and institutions involved
Project contributors
Products and publications
Publications:
Bonifay, W. (2022). Increasing generalizability via the principle of minimum description length. Behavioral and Brain Sciences, 45, Article E5 2022
Bonifay, W., Cai, L., Falk, C. F., & Preacher, K. J. (2025). Reassessing the fitting propensity of factor models. Psychological Methods.
Bonifay, W., & Depaoli, S. (2023). Model evaluation in the presence of categorical data: Bayesian model checking as an alternative to traditional methods. Prevention Science, 24(3), 467-479.
Bonifay, W., Winter, S. D., Skoblow, H. F., & Watts, A. L. (2024). Good fit is weak evidence of replication: increasing rigor through prior predictive similarity checking. Assessment, 10731911241234118.
Cai, L. Hansen, M., Jeon, M., & Edwards, M. C. (2025). Modeling data from educational assessments. In L. L. Cook & M. J. Pitoniak (Eds.), Educational Measurement (5th Ed.) (pp. 682-760). Oxford, UK: Oxford University Press
Davis-Stober, C. P., Dana, J., Kellen, D., McMullin, S. D., & Bonifay, W. (2024). Better accuracy for better science... through random conclusions. Perspectives on Psychological Science, 19(1), 223-243.
Feng, T., & Cai, L. (2026). Sensemaking of process data from evaluation studies of educational games: An application of crossâclassified item response theory modeling. Journal of Educational Measurement, 63(1), e12396.
Watts, A.L., Greene, A.L., Bonifay, W. et al. (2024). A critical evaluation of the p-factor literature. Nature Reviews Psychology, 3, 108-122.
Winter, S. D., Eddy, C. L., Yang, W., & Bonifay, W. (2025). A tutorial on Bayesian item response theory: An illustration using the Teacher Stress Inventory-Short Form. Journal of School Psychology, 109, 101427.
Questions about this project?
To answer additional questions about this project or provide feedback, please contact the program officer.