Unmanned Aerial Vehicles (UAVs) increasingly operate in missions requiring the simultaneous satisfaction of multiple objectives: reaching task locations, performing the correct service, and preserving sufficient onboard energy for continuous operation. Mission efficiency depends not only on task completion but also on managing the trade-off between service duration and battery recharging. This work proposes a Double Deep Q-Network (DDQN) policy for energy-aware UAV navigation on graph maps. The UAV must first collect the appropriate servicing tool from a depot node and then deliver it to the active failure node. At the same time, it autonomously decides when to interrupt the mission for recharging so as to ensure sufficient battery reserve throughout continuous operations, while minimizing task-servicing duration. The key contribution is an energy-aware reward based on a Dynamic Battery Threshold (DBT) computed from graph shortest-path distances to the nearest charging station, enabling a topology-aware recharge policy that is safer yet less conservative than a per-map tuned safety margin. Extensive Monte Carlo tests on increasingly complex graphs show that the proposed policy achieves a 100% task completion rate with always sufficient final battery to reach a charging node from the task node, while degrading less with map complexity and exhibiting greater robustness to stochastic battery dynamics than a pseudo-optimal baseline.

Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning

Gabriele Gemignani
;
Lorenzo Pollini
2026-01-01

Abstract

Unmanned Aerial Vehicles (UAVs) increasingly operate in missions requiring the simultaneous satisfaction of multiple objectives: reaching task locations, performing the correct service, and preserving sufficient onboard energy for continuous operation. Mission efficiency depends not only on task completion but also on managing the trade-off between service duration and battery recharging. This work proposes a Double Deep Q-Network (DDQN) policy for energy-aware UAV navigation on graph maps. The UAV must first collect the appropriate servicing tool from a depot node and then deliver it to the active failure node. At the same time, it autonomously decides when to interrupt the mission for recharging so as to ensure sufficient battery reserve throughout continuous operations, while minimizing task-servicing duration. The key contribution is an energy-aware reward based on a Dynamic Battery Threshold (DBT) computed from graph shortest-path distances to the nearest charging station, enabling a topology-aware recharge policy that is safer yet less conservative than a per-map tuned safety margin. Extensive Monte Carlo tests on increasingly complex graphs show that the proposed policy achieves a 100% task completion rate with always sufficient final battery to reach a charging node from the task node, while degrading less with map complexity and exhibiting greater robustness to stochastic battery dynamics than a pseudo-optimal baseline.
2026
Gemignani, Gabriele; Pollini, Lorenzo
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1369628
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact