Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
Abstract
OBJECTIVES: Vision-language models are increasingly used to interpret medical and everyday images through consumer chat interfaces, yet their ability to read orientation - the single perceptual operation tested by the tumbling-E acuity optotype - is poorly characterized on the surfaces through which they are actually used.
METHODS: We evaluated four production vision-language models (referred to as Claude, GPT, GROK, and Gemini) through their consumer chat interfaces on a locked set of seven optotype charts: four uniform tumbling-E charts (one per cardinal orientation), two mixed-orientation tumbling-E charts, and one Snellen letter chart as a specificity control.
Each model was run in two reasoning modes (Fast and Thinking) under two prompt variants (with and without an explicit orientation-decoding rule) by up to three operators.
The corpus comprised 920 scoreable trials and 50,420 glyph judgements.
The primary outcome was glyph-level accuracy against the chart's designed orientation, summarized with Wilson 95% confidence intervals.
RESULTS: Accuracy ranged from 43.0% to 97.0% across models on identical charts, and the strongest model depended on reasoning mode (GPT 97.0% in Fast mode; GROK 96.6% in Thinking mode).
Errors were not random but collapsed onto a model-specific attractor direction.
Models were 96-100% internally self-consistent yet ranged widely in accuracy, dissociating reliability from validity.
An answer-key-free ensemble-consensus estimate tracked accuracy closely (r = 0.998).
For one model, consumer-interface accuracy fell 25-27 points below programmatic access, almost entirely on a single orientation.
CONCLUSIONS: A single accuracy figure conceals clinically relevant, orientation-specific failure modes; vision-language models should be evaluated along multiple axes and on the deployment surface before image-interpretation outputs are trusted.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요