Validity, Sparse Holes, and Breadth in Language Generation: Banach Density, Topology, and Geometry
Abstract
Language generation in the limit, rooted in work of Gold and Angluin and revived by Kleinberg and Mullainathan, studies generation under minimal assumptions: an adversary enumerates strings from an unknown target language $K$, and an algorithm must eventually generate unseen strings from $K$. A central question is the tension between validity and breadth: can a generator avoid hallucination while covering the hidden language broadly?
In our previous work, we proved that every countable language class admits the optimal 1/2 lower-asymptotic-density guarantee. This prefix-based measure makes validity and breadth universally compatible. Here we replace prefix averaging by the local requirement of lower Banach density, which examines every long interval or box. We prove that for some countable language collections, every eventually valid generator must leave arbitrarily large sparse holes. Thus a worst-case local form of mode collapse is unavoidable: sparse holes are forced by the requirement of valid generation itself.
This local notion reveals structure hidden by prefix averaging. In one dimension, the problem is intrinsically topological and combinatorial: finite Cantor--Bendixson rank guarantees the optimal lower Banach density 1/2, whereas infinite-rank classes can force density zero. In dimensions $d \geq 2$, Ramsey/discrepancy geometry adds in to force ordinary Banach density zero even for a singleton language class. We introduce filtered lower Banach density to remove this geometric obstruction and prove a $1/2-\epsilon$ guarantee for finite-rank classes. We also establish a dichotomy for $f$-window densities, interpolating between lower asymptotic and lower Banach density.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요