Loglinear modelling of huge contingency tables
arXiv:2603.07288v2 Announce Type: replace
Abstract: Contingency tables are the canonical representation of multivariate categorical data. As the size of the contingency table grows exponentially with the number of variables, even a moderate number of variables, each with a moderate number of levels, results in a huge number of cells, the majority of which remains empty even with a significant amount of data. We propose efficient methods for inferring higher-order loglinear models by performing subsampling on the set of the empty cells. First, we derive the likelihood under a zero-deflated Poisson sampling scheme. This is maximized via an efficient iteratively re-weighted least squares algorithm, leading to consistent and close to efficient estimators. This method works well for moderately sized contingency tables, but runs into computational instability when the number of dimensions grows. By sacrificing some efficiency, we show that nested case-control multinomial sampling combined with a degenerate logistic regression approach is also consistent and can be applied to arbitrarily large contingency tables. We illustrate the method with an analysis of data from the General Social Survey, which consists of $15014$ observations in a $69$-dimensional contingency table with a total of $6.6\times 10^{38}$ cells. ...
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요