Simulation-free extrapolation for misspecified models induced by categorizing an error-prone continuous covariate
Abstract
Epidemiological studies often categorize continuous exposures for interpretation even when the underlying outcome-exposure association is continuous.
The fitted categorical regression is a misspecified model because it replaces the continuous exposure with categories.
With measurement error, categorization also misclassifies latent exposure categories, so the observed regression generally targets different means and contrasts.
Existing estimating-equation and simulation-based extrapolation approaches require, respectively, outcome-model-specific derivations and pseudo-data generation with repeated fitting.
We introduce simulation-free extrapolation (SIMFEX), which estimates misclassification probabilities and latent category proportions from replicates, computes the mean-scale trajectory without pseudo-data or repeated outcome-model fitting, and extrapolates it to the no-misclassification endpoint.
Without additional error-free covariates, this construction applies across known one-to-one links, and the resulting estimator is consistent under stated conditions.
With error-free covariates, the contrast relation remains exact for identity-link additive models when misclassification probabilities and category proportions are covariate-invariant; the nonidentity-link version provides a practical approximation.
Simulations show substantial bias reduction relative to the naive analysis and coverage generally close to nominal.
In the UK Biobank analysis, SIMFEX produced larger estimated high-versus-low fat-intake contrasts than the naive analysis for body mass index and obesity, illustrating how category misclassification can change the magnitude and uncertainty of prespecified contrasts.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요