Tools and evaluation in GAZE
- Viewer tools
The model can request viewer-level operations, including zoom, windowing, contrast and edge detection.
- Retrieval tools
Retrieval tools query PubMed literature and Open-i radiological images.
- Joint evaluation
Structured outputs are scored jointly for diagnosis, localisation and captioning.
Method overview from the public workflow documentation and paper.
Abstract. Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times and consult the literature before writing a report. We introduce GAZE (Grounded Agentic Zero-shot Evaluation), a framework that lets a medical VLM work in this iterative way by calling viewer-level tools (zoom, windowing, contrast, edge detection) and two retrieval tools backed by the U.S. National Library of Medicine (PubMed for medical literature, Open-i for radiological images), with structured outputs validated against a schema and full tool-call traces recorded for auditability. On NOVA, a benchmark of 906 brain MRI cases covering 281 rare neurological conditions, GAZE reaches 58.2 mean average precision (mAP) at intersection-over-union (IoU) 0.3 for lesion localisation and 34.9% Top-1 diagnostic accuracy under a joint protocol that scores captioning, diagnosis, and localisation from the image alone, without task-specific fine-tuning. Before any tool is used, structured prompting and schema-validated outputs already improve over the published Gemini 2.0 Flash baseline (20.2 to 29.4 mAP@0.3), so framework design is itself an experimental variable. Tool use helps rare pathologies disproportionately: the fraction of cases with IoU > 0.3 rises from 17% to 58% for diagnoses with three or fewer examples versus 25% to 68% for common conditions (≥10 cases), with gains tracking engagement (Gemini 3 Flash: Cohen’s d = 0.79, 11.8 tool calls per case; Gemini 2.0 Flash: tools used in 8.2% of cases, no significant benefit). Retrieval ablations additionally reveal a model-dependent trade-off in which gains in diagnosis can coincide with losses in localisation, reinforcing the case for joint evaluation of diagnosis, localisation, and captioning in medical VLMs.
Year 2026Kind conference paperVenue International Conference on Artificial Intelligence in Healthcare (AIiH) 2026
Materials
Viewer tools, schema-validated outputs and recorded tool calls for iterative medical VLM evaluation.
Joint localisation, diagnosis and captioning results on NOVA, including model-dependent retrieval effects.
@inproceedings{alim2026gaze,
author = {Alim, Duaa and Alim, Mogtaba and Chalcroft, Liam},
title = {{GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI}},
booktitle = {International Conference on Artificial Intelligence in Healthcare (AIiH) 2026},
year = {2026},
url = {https://arxiv.org/abs/2605.00876}
}