RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
Abstract
Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks.
However, their capability for rare scenes remains insufficiently understood, because existing benchmarks are dominated by common urban and rural imagery.
To address this gap, we present RRS-10K, a benchmark for rare remote sensing image interpretation.
RRS-10K contains 10,738 military-related remote sensing images and corresponding multiple format question-answer pairs for comprehensive evaluation.
All of the images are collected from first-hand sources and organized into three capability dimensions, six sub-dimensions, and 20 leaf tasks, covering perception, reasoning, and robustness.
To improve the quality of multiple-choice questions, we introduce a similarity-based distractor filtering strategy (SDFS) during benchmark construction.
We further evaluate 52 representative models and show that current VLMs achieve only moderate zero-shot performance on rare remote sensing image interpretation, with clear weaknesses in visual grounding, referring segmentation, and complex semantic reasoning tasks.
RRS-10K enables systematic analysis of failure modes in long-tail remote sensing interpretation and provides guidance for developing more reliable remote sensing VLMs.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요
공식 발표 ↔ 진영별 보도
보도 없음
보도 없음
보도 없음