Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
Abstract
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.
Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.
This paper introduces a new method to address this urgent need.
Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores.
The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure.
Our method is simple to implement and does not require token-level labeled data for training.
Theoretically, we characterize this trade-off and show that the proposed method achieves favorable mean square error performance in estimating the underlying signal.
Empirically, we demonstrate strong performance of our method against a wide range of baselines in both synthetic datasets and a realistic dataset.
We deploy a publicly accessible website that implements the methods as well.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요