Skip to main content

Breadcrumb

Home arrow_forward_ios Information on ... arrow_forward_ios Statistical Too ...
Home arrow_forward_ios ... arrow_forward_ios Statistical Too ...
Information on ...
Grant Open

Statistical Tools for Generalizing in Education Research

NCER
Program: Statistical and Research Methodology in Education
Award amount: $897,426
Principal investigator: Andrew Gelman
Awardee:
Columbia University
Year: 2026
Award period: 3 years (08/01/2026 - 07/31/2029)
Project type:
Methodological Innovation
Award number: R305D260042

Purpose

The project aims to develop new statistical tools for for nonrepresentative samples and imbalanced experiments, addressing problems of generalization that are central challenges in empirical research. The project will improve statistics and research methodology in education research by enabling researchers and policymakers to make more reliable generalizations and focused estimates, unifying the ideas of multilevel regression, poststratification, and sampling weights. The goal is for education researchers to be able to fit more sophisticated regressions, multilevel models, and small-area estimation to real-world data.

Structured Abstract

Research design and methods

Currently, multilevel regression and poststratification (MRP) and related methods are state of the art for small-area estimation and adjustment for nonrandom samples in social science, and inverse-probability weighting is the standard design-based method. We propose to combine these approaches by jointly modeling the outcome, predictors, and log weights, an approach which should now be feasible by fitting a simultaneous-equations model using Stan.

The approach will use the following statistical methods: Bayesian inference, sampling theory, bootstrapping, Monte Carlo methods, multilevel regression, and poststratification. The core of the proposed method is a new theoretical idea for embedding sampling weights into a joint distribution with modeled outcomes, along with analytical work mapping population to sampling distribution, and a computational implementation using probabilistic programming. Validation will be performed in three ways: simulated data, artificially simulated subsamples of real data, and real data validated to external benchmarks.

User testing. Education methodologists will be able to experiment with our method and generalize it by extending our software implementation, for example by setting up their own regression models or swapping in machine-learning predictions. Our evaluations will be transparent with full code, so that methodologists can also perform their own evaluations on other datasets.

Applied education researchers will be able to try out the model with their data using scripts that we will provide, following the steps in our case studies, following the model of our multilevel regression and poststratification case study (https://bookdown.org/jl5522/MRP-case-studies/) and Stan case studies for education (e.g., https://mc-stan.org/users/documentation/case-studies/tutorial_twopl.html). The proposal is to develop a new method and also to implement it in a user-friendly and accessible way.

People and institutions involved

IES program contact(s)

Elizabeth Albro

Elizabeth Albro

Commissioner of Special Education Research
NCSER

Project contributors

Tyler W. Watts

Key Personnel

Products and publications

Anticipated statistical and methodological products include: (1) a new method for combining regression (including multilevel modeling and small-area estimation) and inverse-probability modeling using Bayesian inference; (2) validations of this method using real and simulated data; (3) an efficient and user-friendly computational implementation in the open-source probabilistic programming language Stan, with interfaces to R, Python, Stata, Julia, and other general-purpose software; (4) journal publications deriving, explaining, and demonstrating the method; (5) inclusion of these ideas in forthcoming textbooks; (5) tutorial materials and examples specifically focused on education research, including applications to widely-used education datasets (new or improved methods, toolkit, software, guidelines, compendia, review papers, articles) the research team will develop.

Dissemination is essential to this project, as the goal is to develop methods that will be relevant to generalization for item-response models, surveys, experiments, and observational studies that arise in education research in the presence of nonrepresentative samples. As such, we plan to develop full case studies demonstrating the use of these methods, and free and open discussion regarding these tools and case studies will also be available through the online forums for Stan users and developers. We also plan to disseminate the findings through peer-reviewed journal manuscripts, conference presentations, and workshops at major conferences, as well as in forthcoming textbooks and on social media through our widely-read blog. To enable unrestricted use by the largest community of users for education and industry, all of our software is free and open-source. With the IES-supported project Stan, we have a demonstrated track record of disseminating Bayesian methods and software that have been widely adapted by education researchers.

Use in applied education research. Our method and product will be developed and used to improve internal and external validity in an ongoing study of the North Carolina pre-K program. More generally, it will be used in applied education research for analysis of non-representative data for which there are sampling weights. This describes a large number of studies, including the National Assessment of Educational Progress, the Program for International Student Assessment, the National Longitudinal Transition Study, High School and Beyond, Early Childhood Longitudinal Study, and many others. The proposed research will enable researchers who are analyzing such data to integrate weighting into small-area estimation and multilevel regression and poststratification, resulting in more reliable generalizations of inferences to populations of interest.

Related projects

Efficient and Flexible Tools for Complex Multilevel and Latent Variable Modeling in Education Research

R305D190048

Questions about this project?

To answer additional questions about this project or provide feedback, please contact the program officer.

 

Tags

Academic Achievement

Share

Icon to link to Facebook social media siteIcon to link to X social media siteIcon to link to LinkedIn social media siteIcon to copy link value

Questions about this project?

To answer additional questions about this project or provide feedback, please contact the program officer.

 

You may also like

Zoomed in IES logo
Request for Applications

Using Longitudinal Data to Support State Education...

October 01, 2026
Read More
Zoomed in IES logo
Request for Applications

Research Training Programs in the Education Scienc...

October 01, 2026
Read More
Zoomed in IES logo
Fact Sheet/Infographic/FAQ

Understanding the Use of an Early Warning System t...

Author(s): REL Pacific
Read More
icon-dot-govicon-https icon-quote