Shielded RL for Route-Charged Parity-Term Ordering in QEDA Phase Components
Abstract
Commuting phase terms in quantum electronic design automation (QEDA) placement circuits are logically invariant under reordering, yet their routed cost varies substantially after hardware mapping, since term order affects CNOT cancellation, interaction locality, and routing pressure.
We cast parity/support phase-term ordering within a QEDA phase component as a shielded reinforcement-learning problem: a feasibility shield restricts each step to unemitted terms, so every trajectory is a valid permutation by construction, and an elite (cross-entropy-method) policy is trained against a route-charged proxy combining support-transition size and heavy-hex topology-distance features.
We validate by direct Qiskit routing of logically equivalent circuits to a synthetic IBM-style heavy-hex map.
On 36-term parity-walk components (50 term seeds x 2 transpiler seeds, statistics at the term-seed level), the per-instance learned ordering reduces mean routed CX to 336.0, a 5.7-12.2% paired reduction over 2-opt and simulated-annealing search at equal or greater proxy budget and 22.3% over the default construction order; routed-CX and routed-depth gains are significant after Bonferroni correction.
Honest transfer audits show the proxy is predictive for the parity-walk component but not for extraction-heavy or token/permutation circuits, which require architecture-aware rewards, scoping the contribution accordingly.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요