Toward a Unified Statistical Theory of Unsupervised Pretraining and Supervised Neural Knowledge Graph Learning
Abstract
Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications.
However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding.
We address both issues with a theoretically grounded, end-to-end training framework that extends and subsumes existing methods.
Our framework is a two-stage procedure: unsupervised pretraining over heterogeneous corpora followed by supervised learning with multiple relation types.
We establish a nonasymptotic risk bound that disentangles pretraining representation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction.
Synthetic experiments validate each theoretical component, and real-world experiments confirm the effectiveness of our approach on large-scale knowledge graph benchmarks.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요