Predictive Subsampling for Scalable Inference in Networks
Abstract
Current methods for statistical inference in networks often encounter substantial computational bottlenecks when applied to the massive network datasets that are increasingly common across scientific domains.
In this paper, we develop \textit{Predictive Subsampling} (\texttt{PredSub}), a scalable framework for estimation and two-sample testing in networks.
The central idea is to replace a full-sample estimation procedure by estimation on a random subsample, followed by out-of-sample prediction of the remaining vertices.
This construction exploits the fact that the subsample provides an inferential anchor, which enables each remaining vertex to be incorporated in a \textit{predictive} manner through a fast vector operation.
Building on this estimator, we develop two scalable procedures for two-sample testing, namely \texttt{PredSubTest} and \texttt{PureSubTest}.
We establish finite-sample error bounds as well as estimation and testing consistency of the proposed methods in both Frobenius and two-to-infinity norms.
These results formally characterize the trade-offs between statistical accuracy and computational efficiency with respect to the subsample size, the choice of test statistic, and the choice of norm.
We demonstrate the empirical performance of the proposed methods through detailed simulation studies and two real-world applications involving DBLP coauthorship networks and the Cannes 2013 social media networks.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요