Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests
Abstract
Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expressions that reveal governing relationships directly from data.
Existing methods, however, often rely on heuristic search, struggle to balance predictive accuracy with expression complexity in noisy settings, and offer limited characterization of symbolic uncertainty.
Probabilistic approaches that address these challenges in a unified manner remain underexplored.
We introduce a probabilistic symbolic regression framework that represents mathematical expressions as ensembles of symbolic trees.
A regularizing prior over tree topology controls expression complexity, while an Occam's window-based posterior summary captures uncertainty across multiple plausible symbolic models.
Given the limited existing theoretical treatment of symbolic regression, we develop posterior concentration guarantees under approximate symbolic realizability, yielding a near-parametric rate for exact symbolic representability.
Additionally, we establish a sharp oracle concentration result under symbolic misspecification.
Comparisons of our proposed framework with state-of-the-art competitors demonstrate superior predictive accuracy, optimal symbolic complexity, and stable structural recovery when learning benchmark scientific equations, together with the identification of scientifically interpretable descriptors in a challenging materials discovery problem.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요