Perencanaan Lintasan dan Penghindaran Rintangan Multi-UAV Menggunakan Multi-Agent Recurrent Deterministic Policy Gradient (MARDPG)

Yusup, Arya Nugraha Aulia Rahman (2026) Perencanaan Lintasan dan Penghindaran Rintangan Multi-UAV Menggunakan Multi-Agent Recurrent Deterministic Policy Gradient (MARDPG). Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5022221142-Undergraduate_Thesis.pdf] Text
5022221142-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Navigasi multi-unmanned aerial vehicle (multi-UAV) pada lingkungan tiga dimensi yang partially observable menghadapi tantangan berupa non-stationarity, di mana setiap agen hanya memiliki observasi lokal yang terbatas sekaligus harus beradaptasi terhadap perubahan kebijakan agen lain selama proses pembelajaran. Dalam kerangka Decentralized Partially Observable Markov Decision Process (Dec-POMDP), pengambilan keputusan yang optimal memerlukan pemanfaatan riwayat observasi untuk menginferensi keadaan lingkungan yang tidak sepenuhnya dapat diamati. Penelitian ini mengintegrasikan Long Short-Term Memory (LSTM) ke dalam arsitektur Multi-Agent Recurrent Deterministic Policy Gradient (MARDPG) untuk menangkap dependensi temporal pada observasi lokal, serta mengusulkan Curriculum Learning (CL) berbasis difficulty-axis separation guna meningkatkan stabilitas dan efisiensi proses pelatihan melalui peningkatan tingkat kesulitan secara bertahap. Metode yang diusulkan dievaluasi pada skenario navigasi lima UAV di lingkungan tiga dimensi dengan rintangan statis dan dinamis. Hasil evaluasi menunjukkan bahwa CL-MARDPG mencapai success rate sebesar 92,40% dengan collision rate sebesar 6,75%, meningkat dari 87,75% dan menurunkan collision rate dari 11,39% yang diperoleh MARDPG tanpa Curriculum Learning. Dibandingkan dengan MADDPG, metode yang diusulkan meningkatkan success rate dari 74,07% menjadi 92,40% serta menurunkan collision rate dari 25,93% menjadi 6,75%. Hasil tersebut menunjukkan bahwa kombinasi arsitektur recurrent dan Curriculum Learning mampu meningkatkan keandalan navigasi, keselamatan koordinasi, serta stabilitas proses pembelajaran pada lingkungan multi-UAV yang partially observable.
===================================================================================================================================
Multi-unmanned aerial vehicle (multi-UAV) navigation in partially observable three-dimensional environments faces the challenge of non-stationarity, where each agent possesses only limited local observations while simultaneously needing to adapt to changes in other agents' policies during the learning process. Within the Decentralized Partially Observable Markov Decision Process (Dec-POMDP) framework, optimal decision-making requires leveraging observation history to infer environmental states that cannot be fully observed. This study integrates Long Short-Term Memory (LSTM) into the Multi-Agent Recurrent Deterministic Policy Gradient (MARDPG) architecture to capture temporal dependencies in local observations, and proposes a difficulty-axis separation-based Curriculum Learning (CL) approach to enhance training stability and efficiency through a gradual increase in difficulty levels. The proposed method is evaluated on a navigation scenario involving five UAVs in a three-dimensional environment with static and dynamic obstacles. Evaluation results demonstrate that CL-MARDPG achieves a success rate of 92.40% with a collision rate of 6.75%, improving from 87.75% and reducing the collision rate from 11.39% obtained by MARDPG without Curriculum Learning. Compared to MADDPG, the proposed method increases the success rate from 74.07% to 92.40% and decreases the collision rate from 25.93% to 6.75%. These results indicate that the combination of recurrent architecture and Curriculum Learning is capable of improving navigation reliability, coordination safety, and learning process stability in partially observable multi-UAV environments.

Item Type: Thesis (Other)
Uncontrolled Keywords: multi-UAV, MARDPG, Curriculum Learning, Perencanaan Lintasan, Penghindaran Rintangan, multi-UAV, MARDPG, Obstacle Avoidance, Path Planning
Subjects: Q Science > QA Mathematics > QA274.7 Markov processes--Mathematical models.
Q Science > QA Mathematics > QA336 Artificial Intelligence
Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
T Technology > TK Electrical engineering. Electronics Nuclear engineering
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Electrical Engineering > 20201-(S1) Undergraduate Thesis
Depositing User: Arya Nugraha Aulia Rahman Yusup
Date Deposited: 31 Jul 2026 10:03
Last Modified: 31 Jul 2026 10:03
URI: http://repository.its.ac.id/id/eprint/140675

Actions (login required)

View Item View Item