As multimodal interaction becomes central to digital learning, technical interfaces, such as eBooks and conversational agents, present unique accessibility challenges for users relying on the auditory modality. While modern Text-to-Speech (TTS) systems excel at rendering natural language, they often fail to convey the complex, structural syntax of source code, creating a significant barrier for blind and low-vision programmers. We frame the vocal rendering of source code as a cross-modal translation task, in which spatial and symbolic information visually encoded in code must be mapped into a linear auditory signal. This study explores "Code-to-Speech" (CTS) as a critical component of multimodal technical communication. We present a comparative evaluation of five leading AI-driven TTS engines (including ChatGPT, Gemini, and ElevenLabs) to determine their efficacy in rendering mixed-content (text and code) materials. Adopting a mixed-methods approach, we combine automated linguistic metrics (WER, BLEU, METEOR) with a human-centered qualitative assessment. Our findings reveal significant differences: ChatGPT and Gemini consistently achieved lower error rates, while Copilot omitted critical tokens. However, all systems struggle with essential structural cues such as indentation, pauses, and code comments. We also show that designed prompts improve intelligibility, particularly for novices. Our work establishes a benchmark for CTS and underscores the need for multimodal AI architectures that support inclusive technical interaction.

Code-to-Speech: An Exploratory Study of AI Text-to-Speech for Source Code Readability

Marco Cardia;Letizia Angileri;Marina Buzzi;Barbara Leporini
2026-01-01

Abstract

As multimodal interaction becomes central to digital learning, technical interfaces, such as eBooks and conversational agents, present unique accessibility challenges for users relying on the auditory modality. While modern Text-to-Speech (TTS) systems excel at rendering natural language, they often fail to convey the complex, structural syntax of source code, creating a significant barrier for blind and low-vision programmers. We frame the vocal rendering of source code as a cross-modal translation task, in which spatial and symbolic information visually encoded in code must be mapped into a linear auditory signal. This study explores "Code-to-Speech" (CTS) as a critical component of multimodal technical communication. We present a comparative evaluation of five leading AI-driven TTS engines (including ChatGPT, Gemini, and ElevenLabs) to determine their efficacy in rendering mixed-content (text and code) materials. Adopting a mixed-methods approach, we combine automated linguistic metrics (WER, BLEU, METEOR) with a human-centered qualitative assessment. Our findings reveal significant differences: ChatGPT and Gemini consistently achieved lower error rates, while Copilot omitted critical tokens. However, all systems struggle with essential structural cues such as indentation, pauses, and code comments. We also show that designed prompts improve intelligibility, particularly for novices. Our work establishes a benchmark for CTS and underscores the need for multimodal AI architectures that support inclusive technical interaction.
2026
979-8-4007-2318-6
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1366527
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact