KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
Abstract
In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis.
However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines.
Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning.
However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback.
In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification.
Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning.
By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models.
We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback.
We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요