Artificial intelligence has shown promise in dry eye care, but one of the field’s biggest problems has been surprisingly basic: the same image can be labeled different ways by different teams. A new peer-reviewed expert consensus aims to address that by laying out standardized classification, annotation, and quality-control rules for dry eye imaging used in AI applications.
The paper, Expert consensus on classification and annotation methods, processes, and quality control for dry eye imaging in artificial intelligence applications (2025), was published in Intelligent Medicine and listed online by ScienceDirect on December 19, 2025. It is an open-access guideline, not a clinical trial or a product announcement. In plain terms, it does not test a treatment. It defines how researchers should prepare the image data that powers machine-learning models.
That distinction matters. In AI-driven ophthalmology, model performance often depends less on the algorithm itself than on the consistency of the images used to train it. The new consensus is meant to reduce avoidable variation, support multicenter research, and improve the reliability of downstream AI tools for dry eye imaging.
Why Labeling Has Been the Bottleneck
Dry eye disease is not one single image problem. It can be studied through several different imaging approaches, each capturing a different part of the ocular surface or tear film. That variety has helped researchers build AI tools, but it has also created a fragmented data environment.
According to the consensus paper, AI offers new opportunities for dry eye imaging analysis and auxiliary diagnosis, but inconsistent data standards have slowed progress. If one laboratory records a feature one way and another laboratory labels the same feature differently, models trained on those data can become difficult to compare, validate, or reproduce.
That issue is especially important in medical AI, where a model’s apparent accuracy can drop when it is tested on images from a different clinic, camera, or annotation workflow. The new paper is designed to reduce that problem before it starts, by making the labeling process more uniform from the outset.
What the Consensus Standardizes
The guideline focuses on five imaging modalities commonly used in dry eye research and AI work. These are the lipid layer of the tear film, tear meniscus height, tear film breakup time, corneal fluorescein staining, and meibomian gland images.
Each one captures a different clue about tear stability or gland function. The consensus does not just name them; it sets standards for how they should be classified and annotated so different teams can speak the same technical language.
In practice, that means researchers are being given a shared framework for deciding what counts as a usable image, what features should be tagged, and how those tags should be checked for consistency. The paper also describes quality-control requirements such as annotation consistency assessment, multi-round review, and data cleaning.
EurekAlert’s June 3, 2026 release added concrete examples of how that framework may work in practice. Those examples include ghost-gland exclusion, kappa-based consistency checks, and federated-learning or privacy-preserving collaboration. The underlying paper is the source to rely on for the formal consensus, while the release helps illustrate how the recommendations may be applied.
Why the Five Modalities Matter
For readers outside ophthalmology, the value of standardization is easier to understand if the five modalities are broken down in plain language.
Lipid layer of the tear film: This is the oily outer layer of the tear film. If it is unstable or abnormal, tears can evaporate faster. For AI, consistent labeling helps models learn what patterns indicate a healthy or disrupted surface.
Tear meniscus height: This measures the small strip of fluid along the eyelid margin. It is one clue to tear volume. Standardized annotation matters because models need to know whether they are learning a true anatomical measurement or a variation caused by image angle or blur.
Tear film breakup time: This refers to how quickly the tear film starts to break apart after a blink. The consensus is important here because timing-based labels are especially vulnerable to inconsistency if different teams use different definitions or timepoints.
Corneal fluorescein staining: This imaging shows spots or areas on the cornea that take up dye, which can indicate surface damage. Consistent grading is essential because the same staining pattern can be described differently by different annotators.
Meibomian gland images: These show the oil-producing glands in the eyelids. The glands are central to many dry eye cases, and AI studies often depend on correctly identifying gland loss, distortion, or exclusion of artifacts. That is where terms like ghost-gland exclusion become relevant.
In all five cases, the underlying issue is the same: AI models can only learn from what humans label. If the labels are noisy, inconsistent, or incomplete, the model may appear strong in one dataset and fail in another.
What This Means for Multicenter Research
The most practical value of the consensus may be in multicenter studies. Dry eye is common, but no single center can supply enough diverse, high-quality imaging data to build broadly reliable AI systems on its own.
That is why standard annotation rules matter so much. If hospitals and labs can follow the same classification scheme, they can combine data more confidently, compare results across sites, and reduce the amount of cleanup needed later. The paper’s quality-control recommendations are built around that idea.
Annotation consistency assessment, multi-round review, and data cleaning may sound procedural, but they are central to whether a dataset is usable. Kappa-based consistency checks, for example, provide a way to measure how well human annotators agree. If agreement is low, the model may be learning the disagreements rather than the disease signal.
Federated learning also appears in the broader framing. That matters because some institutions may want to collaborate without sharing raw patient images across sites. A standardized annotation system can make privacy-preserving collaboration more realistic, since each site can train or evaluate models using the same definitions even if the data stay local.
What Happens Next
The consensus does not settle the clinical role of dry eye AI. It sets the groundwork for more comparable studies, but it does not prove that standardized labeling will immediately improve diagnosis or patient outcomes. That remains an open question for future research.
A useful comparison is recent evidence on AI for meibomian gland dysfunction, which suggests that performance still has room to improve against human graders. That does not make the new consensus less important. It does mean the field is still at the stage of building reliable inputs before expecting routine clinical deployment.
For labs and hospitals, the near-term task is practical: adopt the annotation rules, build quality-control workflows, and decide whether multicenter collaboration will use shared databases or federated methods. The paper gives researchers a common language for those steps.
The unresolved issue is validation. The next milestone will be studies showing whether datasets built under this consensus actually produce more robust AI models across different imaging systems and patient populations. Until then, the guideline represents an important standardization step, not a finished answer.