Image descriptions are a fundamental aspect of web accessibility, allowing blind and low-vision users to access visual information through assistive technologies such as screen readers. Despite their importance, high-quality alternative text is often missing or inadequate, especially for complex images like diagrams, graphs, and other STEM-related content. Recent advances in generative artificial intelligence and large language models (LLMs) have renewed interest in automatically generating image descriptions, raising the question of whether these systems can reliably support accessibility at scale.In this paper, we present a comparative study of human-written and LLM-generated alternative text for STEM images, focusing on accessibility-critical aspects rather than surface-level textual similarity. Using a curated dataset of images with expert-authored reference descriptions, we evaluate the outputs of state-of-the-art multimodal LLMs through a mixed-methods approach. Our evaluation combines traditional automated metrics, such as BLEU and METEOR, with a human-centered analysis targeting accuracy, informational completeness, and the presence of hallucinations. The results show that while LLMs often produce fluent and seemingly informative descriptions, substantial gaps remain compared to human-written alt text, particularly in conveying formal semantics and essential structural details required for accessibility. We further observe that commonly used automated metrics only partially capture accessibility-relevant errors, underscoring the need for evaluation methodologies based on the needs of screen reader users. We discuss the implications of these findings for the design, evaluation, and deployment of generative AI systems in accessible web and educational contexts.

Human vs LLM-Generated Alt Text for STEM Image Accessibility: A Comparative Study

Marco Cardia
Primo
;
Letizia Angileri;Camilla Poggianti;Barbara Leporini
2026-01-01

Abstract

Image descriptions are a fundamental aspect of web accessibility, allowing blind and low-vision users to access visual information through assistive technologies such as screen readers. Despite their importance, high-quality alternative text is often missing or inadequate, especially for complex images like diagrams, graphs, and other STEM-related content. Recent advances in generative artificial intelligence and large language models (LLMs) have renewed interest in automatically generating image descriptions, raising the question of whether these systems can reliably support accessibility at scale.In this paper, we present a comparative study of human-written and LLM-generated alternative text for STEM images, focusing on accessibility-critical aspects rather than surface-level textual similarity. Using a curated dataset of images with expert-authored reference descriptions, we evaluate the outputs of state-of-the-art multimodal LLMs through a mixed-methods approach. Our evaluation combines traditional automated metrics, such as BLEU and METEOR, with a human-centered analysis targeting accuracy, informational completeness, and the presence of hallucinations. The results show that while LLMs often produce fluent and seemingly informative descriptions, substantial gaps remain compared to human-written alt text, particularly in conveying formal semantics and essential structural details required for accessibility. We further observe that commonly used automated metrics only partially capture accessibility-relevant errors, underscoring the need for evaluation methodologies based on the needs of screen reader users. We discuss the implications of these findings for the design, evaluation, and deployment of generative AI systems in accessible web and educational contexts.
2026
979-8-4007-2372-8
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1365227
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact