Comparing AI Models for Diabetic Retinopathy Detection
Comparing AI Models for Diabetic Retinopathy Detection
August 27, 2026
August 27, 2026
in a darkened room in Tanzania, three people sit in a row covering their eyes. One woman has an eye patch

Dilation for DR screening in Tanzania Photograph: Hugh Bassett

Researchers have independently compared four commercially available artificial intelligence systems for diabetic retinopathy screening using the same set of retinal images from Tanzania, providing rare head-to-head evidence to support decisions about which technologies may be appropriate for different screening programmes.

Published in Diabetes Care, the study evaluated Medios AI/Remidio, MONA, Ophtai and SELENA+ against the same reference-standard human grading. The authors say comparative studies that openly identify commercial AI systems remain uncommon, making it difficult for health services to assess different products before procurement or implementation.

This is particularly important in diabetic retinopathy (DR), where a growing number of AI tools are commercially available but most published evaluations have examined individual systems in isolation. Before this study, no head-to-head comparison of different commercial AI systems had been conducted in an African population.

The researchers, led by the International Centre for Eye Health, first identified potentially suitable AI systems through a scoping review and expert consultation. Of 26 systems initially identified, four ultimately met the study criteria, agreed to participate and allowed their results to be published openly.

The four systems were tested using 2,068 retinal photographs from 689 people with diabetes attending a regional diabetic retinopathy screening programme in Kilimanjaro, Tanzania. None of the images had been used to develop the algorithms being tested. The reference standard was produced by trained human graders working within the English diabetic eye screening programme, with disagreements arbitrated by a senior grader.

Among the 689 participants, 379 had referable diabetic retinopathy, including 93 with proliferative diabetic retinopathy.

Performance varied between systems. Sensitivity for detecting referable DR ranged from 83.9% for Ophtai to 93.7% for Medios AI/Remidio, while specificity ranged from 70.3% for Medios AI/Remidio to 79.0% for Ophtai. MONA achieved sensitivity of 89.5% and specificity of 76.8%, while SELENA+ achieved sensitivity of 91.8% and specificity of 72.9%.

This illustrates an important trade-off between systems. The system with the highest sensitivity had the lowest specificity, while the one with the highest specificity had the lowest sensitivity. For screening programmes, that balance matters because higher sensitivity reduces the chance of missing disease, while lower specificity can result in more people being referred unnecessarily.

When retinal images judged ungradeable by human graders were excluded, sensitivity increased across all four systems to between 91.2% and 96.6%. Performance for proliferative DR was also consistently high: three systems identified all 93 cases, while SELENA+ missed one.

The study was not designed or statistically powered to identify a single “best” system. Instead, the researchers aimed to provide a descriptive comparison, showing how commercially available products perform when tested independently on exactly the same real-world dataset.

The systems also differed in characteristics that could influence which product is suitable for a particular health service. All four had European regulatory approval, but their camera compatibility, referral thresholds and offline functionality varied. Medios AI/Remidio operated offline as standard, while SELENA+ could also be supplied with offline functionality. MONA and Ophtai indicated that offline use could potentially be offered but was not a standard feature.

The authors argue that these factors should be considered alongside diagnostic accuracy when comparing AI products. Different screening programmes may prioritise sensitivity, specificity, offline operation, camera compatibility or other features differently depending on their setting and available resources.

The study also highlights the importance of independent and transparent comparative evaluation. Several commercially available systems could not be included because developers declined to participate or did not respond. The authors argue that open reporting of head-to-head evaluations can provide policymakers and health services with stronger evidence when considering commercial AI products.

The findings show that commercially available AI tools can differ meaningfully in both diagnostic performance and practical characteristics. Comparing these differences on representative datasets can help programmes make more informed choices about which technologies are appropriate for their own screening pathways.

Publication

Cleland C, Bascaran C, Makupa W, Shilio B … Bastawrous A, Macleod D, Burton MJ. Head-to-Head Comparative Evaluation of Four Commercially Available Artificial Intelligence Systems for Detecting Referable Diabetic Retinopathy in a Tanzanian Population. Diabetes Care. August 2026. https://doi.org/10.2337/dc26-0572