Graph Neural Networks1 (GNNs) have become increasingly influential in the analysis of brain graphs, offering a flexible framework for modeling complex patterns of connectivity. Their ability to integrate topological information with node and edge level features has made them particularly attractive for studying neurological and psychiatric disorders such as Alzheimer’s disease, Parkinson’s disease, and Autism Spectrum Disorder (ASD)2. Through widely used parcellation techniques, the brain can be divided into anatomically meaningful regions, which are then represented as nodes in a graph3. This graph- based perspective allows to incorporate multiparametric information, such as morphological measurements from structural MRI or signals from resting-state fMRI, within a unified model, ultimately improving the depth and expressiveness of the learned representations. Despite these advantages, many current applications of GNNs in neuroscience still rely on models that function largely as black boxes. While such models may achieve strong predictive performance, they often offer limited insight into which brain regions or connections drive the final decisions, creating challenges for clinical translation and biological interpretation. Addressing this limitation is essential, especially in conditions like ASD, where altered connectivity patterns are heterogeneous and not yet fully understood. In this work, we propose an interpretable GNN-based framework specifically designed to study the human connectome in the context of ASD. Our approach embeds self-attention mechanisms within the model, enabling the extraction of importance scores that highlight which regions of interest contribute most to the classification task. By leveraging these attention scores, the model not only performs graph-level prediction but also provides interpretable markers that support neuroscientific understanding and foster greater transparency in the decision-making process. Methods: We used the publicly available ABIDE dataset (https://fcon_1000.projects.nitrc.org/indi/abide/), which includes more than 2,000 participants. To reduce heterogeneity, only male subjects aged 5–35 with open-eyes resting-state fMRI scans were selected. From each subject, morphological features from structural MRI (sMRI) and time-series from resting-state fMRI (rs-fMRI) were extracted using the Destrieux atlas. These features formed the basis of graph-structured brain representations in which nodes corresponded to anatomical regions and edges reflected functional connectivity computed via Pearson correlation matrices. The problem was framed as an atlas-based binary whole-graph classification task to distinguish ASD from typically developing (TD) individuals. Node-level relevance was later assessed through self-attention scores. To perform this analysis, we developed a model SAGPooling-GraphSAGE and examined the effect of two parameters with connectomic significance: the edge threshold, controlling graph sparsity, and the pooling ratio, determining how many nodes are retained for classification. A nested 5-fold cross-validation scheme was adopted, with 85% of the data used for training and 15% for testing. The model performance was evaluated using AUC, accuracy, F1 score, precision and recall. The training was conducted using the Binary Cross Entropy loss function and the Stochastic Gradient Descent optimizer. Explainability was a key component of the study. By analyzing the self-attention scores extracted from the pooling layer, we identified the most informative brain regions for each subject and for the group as a whole. Setting the pooling ratio to 100% allowed retrieving the attention values for every node. After subject-wise z-score normalization, globally relevant regions of interest were determined using percentile thresholds (90th, 95th, 99th), providing interpretable insights into both the model decision-making process and patterns associated with ASD. Results: We evaluated a GNN-based framework for classifying ASD versus TD brain graphs, focusing on how graph sparsity affects model performance. By varying edge thresholds and pooling ratios, we found that sparser graphs, generated by higher edge thresholds, produced clearer attention distributions, allowing us to highlight more discriminative brain regions. While pooling ratio influenced node selection, its effect was secondary. We achieved the highest AUC of 72.2 ± 1.8% with thresholds around 50%. In comparison, traditional machine-learning models (Linear Regression, SVM, Decision Trees, Random Forests) trained on the same data achieved AUCs of 57–67%, demonstrating the advantage of graph-based learning in capturing the functional connectome topology. Our results are consistent with ASD vs. TD classification performances reported in literature, indicating robustness despite the dataset heterogeneity. Beyond classification, we used our model to identify key ROIs associated with ASD using self-attention scores distribution. The precuneus, linked to self-referential thinking and social cognition, showed altered connectivity. Additional regions included the posterior dorsal cingulate gyri, angular gyri, and superior temporal gyrus, many of which are part of the Default Mode Network (DMN) and frequently disrupted in ASD. The right angular gyrus, involved in multisensory integration, language, spatial awareness, attention, reasoning, and social cognition, exhibited structural and functional alterations correlated with restrictive and repetitive behaviors and impaired inhibitory control. The posterior cingulate cortex, associated with social-emotional processing and executive functions, showed reduced connectivity in ASD, potentially underlying social- cognitive difficulties and rigidity. The left superior temporal gyrus, crucial for language and social information processing, also showed structural and functional abnormalities. Conclusion: In this work, we present a unified framework for analyzing human connectomes with GNNs and apply it to ASD vs. TD classification using the ABIDE dataset. We integrate structural and functional MRI features into a single graph-based representation, allowing our GNN to achieve competitive classification performance, outperforms traditional machine learning models, and produce interpretable results consistent with established neurobiological markers of ASD. By combining accurate classification with interpretability, we identify brain regions that most strongly influence model decisions. Our methodological contributions include a principled graph construction procedure, multimodal integration of MRI-derived features, and an explainability scheme based on self-attention. While our results are preliminary and not intended for clinical deployment, they demonstrate the potential of graph-based learning in neuroimaging. We also recognize several challenges: the heterogeneity of the ABIDE dataset introduces biases from site effects and demographic imbalances, a common issue in multicentric datasets that are necessary to ensure large sample sizes, statistical power, and generalizability. Nevertheless, this variability can produce site-dependent biases that we must carefully address. Overall, our study highlights the promise of graph-based deep learning for neuroimaging analysis and lays the groundwork for future research in brain graph classification, demonstrating that GNNs combine predictive accuracy with mechanistic insight into brain organization. Acknowledgement: Research partly supported by: Artificial Intelligence in Medicine (AIM) project, funded by INFN-CSN5; PNRR - M4C2 - I1.3, PE00000013 “FAIR - Future Artificial Intelligence Research” - Spoke 8 “Pervasive AI”, and PNRR - M4C2 - I1.4, CN00000013 - “ICSC – Centro Nazionale di Ricerca in High Performance Computing, Big Data and Quantum Computing” - Spoke 8 “In Silico medicine and Omics Data”, both funded by the European Commission under the NextGeneration EU programme; SC was partly supported by the Italian Ministry of Health (Ricerca Corrente Linea-4) and AIMS2-Trials, http://aims-2-trials.eu.
An Interpretable Graph Neural Network approach to study Autism Spectrum Disorder
Giuseppe Antonio Motisi
Primo
;Francesca Lizzi;Francesca Mainas;Piernicola OlivaPenultimo
;Alessandra ReticoUltimo
2026-01-01
Abstract
Graph Neural Networks1 (GNNs) have become increasingly influential in the analysis of brain graphs, offering a flexible framework for modeling complex patterns of connectivity. Their ability to integrate topological information with node and edge level features has made them particularly attractive for studying neurological and psychiatric disorders such as Alzheimer’s disease, Parkinson’s disease, and Autism Spectrum Disorder (ASD)2. Through widely used parcellation techniques, the brain can be divided into anatomically meaningful regions, which are then represented as nodes in a graph3. This graph- based perspective allows to incorporate multiparametric information, such as morphological measurements from structural MRI or signals from resting-state fMRI, within a unified model, ultimately improving the depth and expressiveness of the learned representations. Despite these advantages, many current applications of GNNs in neuroscience still rely on models that function largely as black boxes. While such models may achieve strong predictive performance, they often offer limited insight into which brain regions or connections drive the final decisions, creating challenges for clinical translation and biological interpretation. Addressing this limitation is essential, especially in conditions like ASD, where altered connectivity patterns are heterogeneous and not yet fully understood. In this work, we propose an interpretable GNN-based framework specifically designed to study the human connectome in the context of ASD. Our approach embeds self-attention mechanisms within the model, enabling the extraction of importance scores that highlight which regions of interest contribute most to the classification task. By leveraging these attention scores, the model not only performs graph-level prediction but also provides interpretable markers that support neuroscientific understanding and foster greater transparency in the decision-making process. Methods: We used the publicly available ABIDE dataset (https://fcon_1000.projects.nitrc.org/indi/abide/), which includes more than 2,000 participants. To reduce heterogeneity, only male subjects aged 5–35 with open-eyes resting-state fMRI scans were selected. From each subject, morphological features from structural MRI (sMRI) and time-series from resting-state fMRI (rs-fMRI) were extracted using the Destrieux atlas. These features formed the basis of graph-structured brain representations in which nodes corresponded to anatomical regions and edges reflected functional connectivity computed via Pearson correlation matrices. The problem was framed as an atlas-based binary whole-graph classification task to distinguish ASD from typically developing (TD) individuals. Node-level relevance was later assessed through self-attention scores. To perform this analysis, we developed a model SAGPooling-GraphSAGE and examined the effect of two parameters with connectomic significance: the edge threshold, controlling graph sparsity, and the pooling ratio, determining how many nodes are retained for classification. A nested 5-fold cross-validation scheme was adopted, with 85% of the data used for training and 15% for testing. The model performance was evaluated using AUC, accuracy, F1 score, precision and recall. The training was conducted using the Binary Cross Entropy loss function and the Stochastic Gradient Descent optimizer. Explainability was a key component of the study. By analyzing the self-attention scores extracted from the pooling layer, we identified the most informative brain regions for each subject and for the group as a whole. Setting the pooling ratio to 100% allowed retrieving the attention values for every node. After subject-wise z-score normalization, globally relevant regions of interest were determined using percentile thresholds (90th, 95th, 99th), providing interpretable insights into both the model decision-making process and patterns associated with ASD. Results: We evaluated a GNN-based framework for classifying ASD versus TD brain graphs, focusing on how graph sparsity affects model performance. By varying edge thresholds and pooling ratios, we found that sparser graphs, generated by higher edge thresholds, produced clearer attention distributions, allowing us to highlight more discriminative brain regions. While pooling ratio influenced node selection, its effect was secondary. We achieved the highest AUC of 72.2 ± 1.8% with thresholds around 50%. In comparison, traditional machine-learning models (Linear Regression, SVM, Decision Trees, Random Forests) trained on the same data achieved AUCs of 57–67%, demonstrating the advantage of graph-based learning in capturing the functional connectome topology. Our results are consistent with ASD vs. TD classification performances reported in literature, indicating robustness despite the dataset heterogeneity. Beyond classification, we used our model to identify key ROIs associated with ASD using self-attention scores distribution. The precuneus, linked to self-referential thinking and social cognition, showed altered connectivity. Additional regions included the posterior dorsal cingulate gyri, angular gyri, and superior temporal gyrus, many of which are part of the Default Mode Network (DMN) and frequently disrupted in ASD. The right angular gyrus, involved in multisensory integration, language, spatial awareness, attention, reasoning, and social cognition, exhibited structural and functional alterations correlated with restrictive and repetitive behaviors and impaired inhibitory control. The posterior cingulate cortex, associated with social-emotional processing and executive functions, showed reduced connectivity in ASD, potentially underlying social- cognitive difficulties and rigidity. The left superior temporal gyrus, crucial for language and social information processing, also showed structural and functional abnormalities. Conclusion: In this work, we present a unified framework for analyzing human connectomes with GNNs and apply it to ASD vs. TD classification using the ABIDE dataset. We integrate structural and functional MRI features into a single graph-based representation, allowing our GNN to achieve competitive classification performance, outperforms traditional machine learning models, and produce interpretable results consistent with established neurobiological markers of ASD. By combining accurate classification with interpretability, we identify brain regions that most strongly influence model decisions. Our methodological contributions include a principled graph construction procedure, multimodal integration of MRI-derived features, and an explainability scheme based on self-attention. While our results are preliminary and not intended for clinical deployment, they demonstrate the potential of graph-based learning in neuroimaging. We also recognize several challenges: the heterogeneity of the ABIDE dataset introduces biases from site effects and demographic imbalances, a common issue in multicentric datasets that are necessary to ensure large sample sizes, statistical power, and generalizability. Nevertheless, this variability can produce site-dependent biases that we must carefully address. Overall, our study highlights the promise of graph-based deep learning for neuroimaging analysis and lays the groundwork for future research in brain graph classification, demonstrating that GNNs combine predictive accuracy with mechanistic insight into brain organization. Acknowledgement: Research partly supported by: Artificial Intelligence in Medicine (AIM) project, funded by INFN-CSN5; PNRR - M4C2 - I1.3, PE00000013 “FAIR - Future Artificial Intelligence Research” - Spoke 8 “Pervasive AI”, and PNRR - M4C2 - I1.4, CN00000013 - “ICSC – Centro Nazionale di Ricerca in High Performance Computing, Big Data and Quantum Computing” - Spoke 8 “In Silico medicine and Omics Data”, both funded by the European Commission under the NextGeneration EU programme; SC was partly supported by the Italian Ministry of Health (Ricerca Corrente Linea-4) and AIMS2-Trials, http://aims-2-trials.eu.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


