A comprehensive guideline for regularization-path variable selection in high-dimensional Gaussian linear regression
Abstract
This paper provides a comprehensive comparison of complete regularization-path-based variable selection procedures in high-dimensional Gaussian linear regression.
Our simulation study stands out from previous ones for several reasons.
First, it offers comprehensive guidelines for complete variable selection, from regularization path construction to final variable subset selection.
In particular, we compare jointly two convex regularization functions (Lasso and Elastic-Net) with two optimization algorithms (LARS and cyclic coordinate descent) and a range of model selection and variable identification techniques.
Second, we also incorporate methods with non-asymptotic theoretical guarantees, which are typically not included in existing reviews.
It is unrealistic to expect a single method that works best.
Still, we provide detailed guidance on which methods perform best under different evaluation criteria, including pROC-AUC, MSE, recall, specificity, and FDP, and that are robust to data scenarios, including varying variable dependency structures.
Overall, Elastic-Net combined with the LARS algorithm provides the most reliable regularization path, while the preferred final selection procedure depends on whether prediction performances, support recovery or false discovery control is prioritized.
We show that some methods are empirically better for a given criterion than others, even though the latter were designed to control for it theoretically.
Regarding new developments in a non-asymptotic setting, we highlight the quality of LinSelect and the need, as future work, to fine-tune the unknown constants of data-dependent penalties in high-dimensional Gaussian linear regression models.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요