Presented By: Department of Statistics
When Prediction Meets Randomization: Regression, Machine Learning, or LLMs? Lessons from 125 Randomized Trials
Bingkai Wang
Covariate adjustment is a standard tool for improving precision in randomized trials, but the value of increasingly flexible prediction methods remains unclear. This talk begins with an empirical comparison of regression and machine-learning-based adjustment across 50 completed trials. The results show that flexible machine learning methods are not automatically superior: simple regression adjustment with prognostic baseline covariates is often highly competitive. This finding motivates a new question: can large language models change the picture by extracting useful prognostic information from baseline covariates, trial descriptions, or other text-derived features? To address this question, I present a unified framework for LLM-assisted covariate adjustment and empirical evidence from 125 randomized trials. We evaluate LLM-derived features, including zero/few-shot predictions, fine-tuned models, and embedding-based representations, combined with both regression and machine-learning estimators. The results show that LLM features can provide additional precision gains beyond classical covariates, although the gains vary across trials, sample sizes, and outcome types. Overall, LLMs appear most useful as tools for constructing better prognostic features within transparent and honest adjustment procedures.
Department reception in 450 West Hall after the seminar at 11:00 AM
Department reception in 450 West Hall after the seminar at 11:00 AM