MARATTO

article · Procedia Computer Science

Dynamic Reward-Based Deep Reinforcement Learning Algorithm for UAV Path Planning in Large-Scale Environments

20251 citationOpen accessUniversity of Tunis El Manar

Abstract

Path planning for Unmanned Aerial Vehicles (UAV) is a vital component of navigation in robotics. The reinforcement Q-learning algorithm enhances path planning for drones but suffers from the need for a large Q-value table and challenges in complex navigation situations. By integrating deep learning with reinforcement one, these shortcomings can be addressed. In this paper, a Deep Q-Network (DQN) model is developed and trained to estimate the drone’s state-action value function. In this work, the flight space is represented by a grid of cells, which are then encoded to convert environmental information into a new input format suitable for the DQN model. Normalizing state inputs enhances the stability and convergence of the proposed DQN algorithm by ensuring comparability of features across different scales. Besides, a new dynamic reward function is established based on the distance between the drone’s current position and its destination. Simulation results and discussion illustrate the effectiveness of the proposed DQN-based approach for collision-free path planning of UAV in complex environments.

Research topics

  • Robotic Path Planning Algorithms
  • UAV Applications and Optimization
  • Reinforcement Learning in Robotics

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1016/j.procs.2025.09.189

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.