Explaining Siamese Networks in Few-Shot Learning for Audio Data

Fedele, A.; Guidotti, R.; Pedreschi, D.

doi:10.1007/978-3-031-18840-4_36

Machine learning models are not able to generalize correctly when queried on samples belonging to class distributions that were never seen during training. This is a critical issue, since real world applications might need to quickly adapt without the necessity of re-training. To overcome these limitations, few-shot learning frameworks have been proposed and their applicability has been studied widely for computer vision tasks. Siamese Networks learn pairs similarity in form of a metric that can be easily extended on new unseen classes. Unfortunately, the downside of such systems is the lack of explainability. We propose a method to explain the outcomes of Siamese Networks in the context of few-shot learning for audio data. This objective is pursued through a local perturbation-based approach that evaluates segments-weighted-average contributions to the final outcome considering the interplay between different areas of the audio spectrogram. Qualitative and quantitative results demonstrate that our method is able to show common intra-class characteristics and erroneous reliance on silent sections.

Explaining Siamese Networks in Few-Shot Learning for Audio Data

Fedele A.^Primo;Guidotti R.^Secondo;Pedreschi D.^Ultimo

2022-01-01

Abstract

Machine learning models are not able to generalize correctly when queried on samples belonging to class distributions that were never seen during training. This is a critical issue, since real world applications might need to quickly adapt without the necessity of re-training. To overcome these limitations, few-shot learning frameworks have been proposed and their applicability has been studied widely for computer vision tasks. Siamese Networks learn pairs similarity in form of a metric that can be easily extended on new unseen classes. Unfortunately, the downside of such systems is the lack of explainability. We propose a method to explain the outcomes of Siamese Networks in the context of few-shot learning for audio data. This objective is pursued through a local perturbation-based approach that evaluates segments-weighted-average contributions to the final outcome considering the interplay between different areas of the audio spectrogram. Qualitative and quantitative results demonstrate that our method is able to show common intra-class characteristics and erroneous reliance on silent sections.

Scheda breve

Scheda completa

Scheda completa (DC)

Anno

2022

Codice ISBN

978-3-031-18839-8
978-3-031-18840-4

File in questo prodotto:

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1162771

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

5

5

CINECA IRIS Institutional Research Information System

Explaining Siamese Networks in Few-Shot Learning for Audio Data

Fedele A.^Primo;Guidotti R.^Secondo;Pedreschi D.^Ultimo

Primo

Secondo

Ultimo

2022-01-01

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

Attenzione

Citazioni

social impact

CINECA IRIS Institutional Research Information System

Explaining Siamese Networks in Few-Shot Learning for Audio Data

Fedele A.Primo;Guidotti R. Secondo;Pedreschi D.Ultimo

Primo

Secondo

Ultimo

2022-01-01

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Informazioni

Attenzione

Citazioni

social impact

Conferma cancellazione

Fedele A.^Primo;Guidotti R.^Secondo;Pedreschi D.^Ultimo

Scheda breve

Scheda completa

Scheda completa (DC)