As multimodal interaction becomes central to digital learning, technical interfaces, such as eBooks and conversational agents, present unique accessibility challenges for users relying on the auditory modality. While modern Text-to-Speech (TTS) systems excel at rendering natural language, they often fail to convey the complex, structural syntax of source code, creating a significant barrier for blind and low-vision programmers. We frame the vocal rendering of source code as a cross-modal translation task, in which spatial and symbolic information visually encoded in code must be mapped into a linear auditory signal. This study explores "Code-to-Speech" (CTS) as a critical component of multimodal technical communication. We present a comparative evaluation of five leading AI-driven TTS engines (including ChatGPT, Gemini, and ElevenLabs) to determine their efficacy in rendering mixed-content (text and code) materials. Adopting a mixed-methods approach, we combine automated linguistic metrics (WER, BLEU, METEOR) with a human-centered qualitative assessment. Our findings reveal significant differences: ChatGPT and Gemini consistently achieved lower error rates, while Copilot omitted critical tokens. However, all systems struggle with essential structural cues such as indentation, pauses, and code comments. We also show that designed prompts improve intelligibility, particularly for novices. Our work establishes a benchmark for CTS and underscores the need for multimodal AI architectures that support inclusive technical interaction.
Code-to-Speech: An Exploratory Study of AI Text-to-Speech for Source Code Readability
Marco Cardia;Letizia Angileri;Marina Buzzi;Barbara Leporini
2026-01-01
Abstract
As multimodal interaction becomes central to digital learning, technical interfaces, such as eBooks and conversational agents, present unique accessibility challenges for users relying on the auditory modality. While modern Text-to-Speech (TTS) systems excel at rendering natural language, they often fail to convey the complex, structural syntax of source code, creating a significant barrier for blind and low-vision programmers. We frame the vocal rendering of source code as a cross-modal translation task, in which spatial and symbolic information visually encoded in code must be mapped into a linear auditory signal. This study explores "Code-to-Speech" (CTS) as a critical component of multimodal technical communication. We present a comparative evaluation of five leading AI-driven TTS engines (including ChatGPT, Gemini, and ElevenLabs) to determine their efficacy in rendering mixed-content (text and code) materials. Adopting a mixed-methods approach, we combine automated linguistic metrics (WER, BLEU, METEOR) with a human-centered qualitative assessment. Our findings reveal significant differences: ChatGPT and Gemini consistently achieved lower error rates, while Copilot omitted critical tokens. However, all systems struggle with essential structural cues such as indentation, pauses, and code comments. We also show that designed prompts improve intelligibility, particularly for novices. Our work establishes a benchmark for CTS and underscores the need for multimodal AI architectures that support inclusive technical interaction.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


