MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback
Abstract
Causal effect estimation from observational data requires careful adjustment for confounding.
Classical estimators such as inverse probability weighting and augmented inverse probability weighting can perform well under favorable model specification but may become unstable in complex settings.
Machine-learning and representation-learning methods provide greater flexibility, but joint optimization may allow outcome information to alter treatment representations and compromise the intended causal structure.
We propose MOCA (Modular One-way Causal Attention), a transformer-based framework that separates treatment and outcome modeling through a modular architecture.
This design preserves directional information flow while retaining the flexibility of transformer architectures.
Using a Markov-kernel formulation, we show that MOCA learns an autonomous treatment representation and is predictively KL-optimal within its Gaussian model class.
Under correct specification, MOCA recovers the true average treatment effect, while additional two-way feedback does not further reduce average treatment effect estimation error.
We also propose a conformal inference procedure for individual treatment effects.
Across multiple simulation scenarios, MOCA achieved competitive or improved average treatment effect estimation compared with IPW, AIPW, the X-learner, TARNet, DragonNet, BART, Causal Forest, and Do-PFN.
Ablation studies supported the contributions of the proposed architectural components.
We further evaluated MOCA on the Infant Health and Development Program benchmark and the observational Dehejia-Wahba dataset.
Overall, modular attention with one-way information flow provides an effective and interpretable framework for causal inference using modern deep-learning models.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요