Short Research Paper Accepted at the Top International Conference ACM International Conference on Information and Knowledge Management (CIKM) 2026
Files

The paper 'Grounded Distractor Probing: A Diagnostic Protocol for Hard-Distractor Sensitivity in VLM Multiple-Choice Evaluation', written by Oh Seung-hoon (master's, first author), Jung Kang-hyun (master's) and Professor Park Un-sang (corresponding author) of the Computer Vision and Image Processing Lab (CVIP), has been accepted for publication at the top international conference ACM International Conference on Information and Knowledge Management (CIKM) 2026 (acceptance rate: 236/764 = 30.9%).
Vision-language models (VLMs) are now used in a wide range of areas including image question answering, document understanding, chart analysis and scene reasoning. Multiple-choice evaluation, in which the model picks one of several answer candidates, is widely used to measure their performance, but the fact that a model gets the answer right is not enough to conclude that it accurately understood the details of the image.
In a multiple-choice question, a model can pick the right answer by relying on the surface form of the options, the position of the correct answer, linguistic bias or coarse image cues, without closely checking the key evidence in the image. When the existing distractors differ greatly from the correct answer in type or phrasing, the model can eliminate them easily without making full use of the actual visual information.
With this in mind, the team proposes the Grounded Distractor Probing (GDP) protocol, which uses distractors of the same type as the correct answer that can only be ruled out by checking fine visual details such as small numbers in the image, the position of clock hands or the number of objects. GDP compares three conditions on the same image and question: first, the original condition, which uses the existing question as it is; second, a permutation-control condition, which keeps the meaning of the options but changes only their order; and third, a hard-distractor condition, which replaces one existing distractor with a visually plausible near-miss distractor of the same type as the correct answer. This design makes it possible to distinguish whether a drop in performance is caused simply by a change in option order or by hard distractors that require fine visual grounding. The team also compared a direct approach, in which the model picks an answer immediately, with an evidence-first approach, in which it first states the evidence found in the image and then picks an answer.
The team selected items corresponding to TextVQA, InfoVQA, DocVQA, ChartQA and RealWorldQA from the VMCBench development data and built GDP-Slice by constructing, for each item, a near-miss hard distractor of the same type as the correct answer. To check whether the main phenomena were confined to a small diagnostic dataset, they also carried out an additional scale check on GDP-v2-lite, produced through automatic proposal and filtering. GDP-v2-lite is not an independent benchmark replacing GDP-Slice but a supplementary analysis used to confirm that the main diagnostic patterns also appear at larger scale. Rather than presenting a new all-purpose prompt, the research is significant in proposing a diagnostic protocol for analysing how vision-language models respond to distractors that require fine visual grounding.
The results show that visual understanding should not be judged from overall accuracy alone in VLM multiple-choice evaluation; which hard distractors a model is vulnerable to, and what effect reasoning interventions such as stating evidence have on each model, must be measured as well. This allows a more precise analysis in real applications not only of the model's answers but of the stability and fragility of the process by which it arrives at them.
Oh Seung-hoon, the paper's first author and a master's student, said: "This research began from the problem that a vision-language model getting the answer right is not enough to judge that it really used the fine detail of the image. We found that when presented with distractors of the same type as the correct answer that can only be ruled out by checking small differences in the image, having the model state its evidence first helps some models but can increase confusion in others. Going forward I want to extend evaluation research that analyses not only accuracy but which distractors a model is vulnerable to and whether intervening in the reasoning process actually helps. I am grateful to my co-researcher Jung Kang-hyun and to Professor Park Un-sang for his guidance."
The ACM International Conference on Information and Knowledge Management (CIKM) is a globally prestigious international conference covering information retrieval, knowledge management, data mining, databases and artificial intelligence. A CIKM short paper corresponds to a recognized IF of 2 under the BK21 evaluation criteria for outstanding international conferences in computer science, and CIKM is classified as a top conference in the Korean Institute of Information Scientists and Engineers' list of outstanding conferences in the software field. This year it will be held in Rome, Italy, from 7 to 11 November.
References:
● 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)
● Website: https://cikm2026.diag.uniroma1.it/