ADP 前沿学习

← 板块一 · 研究前沿

Valid inference for regression with best subset selection

Lin, Tan, Li · stat.ME · 2026-09-05 · 原文

Best subset selection is widely implemented in statistical software and is routinely used by practitioners in scientific fields for variable selection. However, classical confidence intervals and p-values constructed after model selection generally fail to achieve their nominal frequentist guarantees, which can invalidate subsequent findings. In this article, we characterize the altered conditional sampling distributions of pivotal quantities after best subset selection. Building on selective-inference techniques developed in related settings, our finite-sample characterization of the AIC selection event reveals that its geometry is a union of finitely many intervals on the real line. This geometry enables exact conditioning and avoids the excessive conditioning common in other post-selection methods. We use this characterization to develop valid inference procedures that provide p-values with nominal Type~I error and confidence intervals with finite-sample coverage guarantees. The proposed methods are easy to implement, computationally efficient, and broadly applicable to other commonly used best subset selection criteria. We also study inference with unknown noise level using

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

Best subset selection is widely implemented in statistical software and is routinely used by practitioners in scientific fields for variable selection.

However, classical confidence intervals and p-values constructed after model selection generally fail to achieve their nominal frequentist guarantees, which can invalidate subsequent findings.

2. 领域脉络

本文类目:stat.ME,属于其所在研究脉络的最新进展。

3. 机制拆解

Building on selective-inference techniques developed in related settings, our finite-sample characterization of the AIC selection event reveals that its geometry is a union of finitely many intervals on the real line.

This geometry enables exact conditioning and avoids the excessive conditioning common in other post-selection methods.

4. 证据与数字

摘要未给出量化结果——留意原文的实验与数据。

5. 反例与边界

In this article, we characterize the altered conditional sampling distributions of pivotal quantities after best subset selection.

The proposed methods are easy to implement, computationally efficient, and broadly applicable to other commonly used best subset selection criteria.

6. 跨领域连接与意外收获

思考本文机制能否迁移到你正在跟进的问题。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。