Ilya Pershin
Ilya Pershin
Email
i.pershin (at)
innopolis.ru
Lab
Med AI Lab,
Innopolis University

Research

Human attention is a signal that is hard to fake and almost never used. An expert’s eyes move differently when they are tired, when they are looking at something unfamiliar, and shortly before they make a mistake. Those same movements carry spatial information that would otherwise take hours of manual annotation to produce. My work runs along three lines: analysing that signal, using it, and generating it.

Analysing expert gaze

The question. Diagnostic quality is normally assessed after the fact, from the report. By then the error has been made. Eye movements are recorded continuously and in real time — can they say something about the state of the clinician while they work?

We have shown that a radiologist’s gaze changes measurably with accumulated workload, that it changes according to the type of abnormality on the image, and — the strongest claim — that an impending diagnostic error can be predicted from the gaze that precedes it. Gaze is therefore a candidate instrument for two things radiology currently lacks: an objective measure of reader fatigue, and an early warning of error before it reaches the report.

Diagnostic accuracy of four radiologists across a shift, measured every 100 X-rays. Three of the four recover after the lunch break (200→300) and decline over the last hundred images. From IEEE Journal of Biomedical and Health Informatics, 2022.

Gaze as a control signal for interactive segmentation

The question. Annotating medical images is slow and expensive, and it matters not only for building datasets but for radiotherapy planning. A radiologist reading an image is already producing localisation data — they look at what matters. Can expert eye movements be used to drive interactive segmentation?

Gaze is used for correction: the clinician looks over the wrongly segmented regions and the prediction updates — which is faster than redrawing a contour by hand.

Liver and spleen masks during gaze correction and at the end of it: reference contour in green, model prediction in yellow, gaze points in blue. From IEEE Access, 2025 (Q1).

Modelling human attention

The question. One of the central obstacles in eye tracking is that gaze data is hard to collect: it takes experiments with human participants. It is made worse by the fact that every task needs its own data. One way out is to generate synthetic gaze.

We model human scanpaths during reading, transfer those models across languages, and use the synthetic trajectories as an additional signal in RLHF.

A real and a model-generated scanpath over the same sentence. Arrows are eye movements, colour intensity is average fixation count. From Multilingual Synthetic Scanpaths, ICML 2026 workshop.

Where the lab is going

Three further lines are active now:

Non-invasive risk profiling for neurodegenerative disease

Electronic health record analysis

IVF embryo assessment