Action Quality Assessment (AQA) plays an important role in evalu- ating human performance in different domains, including fitness, sports, and healthcare. This work introduces a novel AQA approach by fine-tuning large multimodal models (LMMs) for personalized activity evaluation. We used the Fitness-AQA Dataset, which pro- vides detailed annotations of exercise errors under realistic con- ditions, and we adapt the LLaVA-Video model, a state-of-the-art LMM comprising the Qwen2 large language model and the SigLIP vision encoder. We have implemented a customized data prepara- tion pipeline that transforms video-based exercise annotations into a conversational format specific for fine-tuning. To our knowledge, this study is among the first to fine-tune LMMs for AQA tasks and the very first to explore activity evaluation in this context. The experimental evaluation shows that our model achieves results slightly lower than the baseline, even though it is able to generalize across multiple exercises. The full-reproducible code is available on GitHub

Fine-Tuning Large Multimodal Models for Fitness Action

Elio Musacchio;
2025-01-01

Abstract

Action Quality Assessment (AQA) plays an important role in evalu- ating human performance in different domains, including fitness, sports, and healthcare. This work introduces a novel AQA approach by fine-tuning large multimodal models (LMMs) for personalized activity evaluation. We used the Fitness-AQA Dataset, which pro- vides detailed annotations of exercise errors under realistic con- ditions, and we adapt the LLaVA-Video model, a state-of-the-art LMM comprising the Qwen2 large language model and the SigLIP vision encoder. We have implemented a customized data prepara- tion pipeline that transforms video-based exercise annotations into a conversational format specific for fine-tuning. To our knowledge, this study is among the first to fine-tune LMMs for AQA tasks and the very first to explore activity evaluation in this context. The experimental evaluation shows that our model achieves results slightly lower than the baseline, even though it is able to generalize across multiple exercises. The full-reproducible code is available on GitHub
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1322968
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact