OPOD: On-Policy Omni Distillation
Abstract
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult.
Training a single model on pooled multimodal data often fails to match models specialized for individual modalities.
On-policy distillation (OPD) offers a way to combine such specialists: the student generates a response, and a teacher evaluates that same response, so the student learns directly from behaviors it actually produces.
Yet using several teachers can introduce competing guidance and improve one modality at the expense of another.
We present On-Policy Omni Distillation (OPOD), which routes each student response to the matching text, image, or audio teacher.
OPOD keeps teacher guidance only on tokens where the teacher assigns a higher probability than the student, adjusts the influence of each modality teacher independently during training, and asks the routed teacher to assess both the final answer and whether the reasoning supports the correct answer.
Across twelve benchmarks and three backbone sizes, OPOD achieves the best average score at every scale, reaching 70.8, 51.7, and 46.2 and exceeding the strongest comparator by 2.1, 1.8, and 1.7 points.
On the 30B model, it outperforms both the base model and a counterpart post-trained jointly on pooled multimodal data on all twelve benchmarks, and ranks first or second on eleven even when the individual specialists are included.
The specialists are discarded after training, leaving one deployable omni-modal model.
These results show that coordinating modality-specific teachers is an effective way to improve a shared model while maintaining cross-modal balance.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요