Arsitektur Terdistribusi untuk Kolaborasi Manusia-Robot Asistif Berbasis Pemrosesan Bahasa Alami dan Kecerdasan Ambien

Muhtadin, Muhtadin (2026) Arsitektur Terdistribusi untuk Kolaborasi Manusia-Robot Asistif Berbasis Pemrosesan Bahasa Alami dan Kecerdasan Ambien. Doctoral thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 07111960010019-Doctoral.pdf] Text
07111960010019-Doctoral.pdf - Accepted Version
Restricted to Repository staff only

Download (26MB) | Request a copy

Abstract

Peningkatan populasi lansia menuntut pengembangan robot layanan yang mampu berkolaborasi secara alami melalui bahasa. Namun, beban komputasi pemrosesan bahasa alami berbasis Large Language Model (LLM) melampaui kapasitas perangkat edge pada robot, sehingga lapisan penalaran harus dipisahkan secara fisis ke server. Pemisahan tersebut menimbulkan persoalan turunan yang justru menjadi inti disertasi ini: lapisan penalaran yang dijauhkan dari robot kehilangan akses terhadap keadaan dunia nyata, sementara instruksi pengguna lansia umumnya tidak lengkap dan sangat bergantung pada konteks. Disertasi ini mengajukan tesis bahwa kecerdasan ambien bukan sekadar fitur tambahan, melainkan prasyarat arsitektural yang memungkinkan lapisan penalaran didistribusikan tanpa kehilangan grounding. Arsitektur tersebut terdiri atas tiga level: Level 1 (kecerdasan ambien pada lapisan edge, mencakup lokalisasi Ultrawideband dan persepsi gestural) yang memberikan informasi konteks; Level 2 (penalaran kognitif pada server, menggunakan Gemma-2B yang di-fine-tune melalui QLoRA, dilengkapi mekanisme validasi human-inthe- loop) yang memetakan perintah natural dan konteks menjadi rencana terstruktur; serta Level 3 (eksekusi fisis kembali pada lapisan edge, mencakup navigasi dan manipulasi) yang menerjemahkan rencana simbolik menjadi aksi motorik. Validasi ditempuh melalui strategi kombinasi bertahap guna melokalisasi sumber galat pada tiap level. Temuan paling menentukan diperoleh dari perbandingan antarmodel yang dijalankan dengan prompt berkonstrain skema yang sama. Pada perintah langsung yang tidak menuntut grounding spasial, kedua model dasar berkapasitas jauh lebih besar unggul, yaitu GPT-3.5 mencapai 97% dan Gemma-7B mencapai 94%, sedangkan Gemma-2B hasil training fine-tune hanya mencapai 89%. Keunggulan tersebut berbalik pada tugas penalaran spasial, yaitu Gemma-2B bertahan pada 87% sementara GPT-3.5 turun menjadi 15% dan Gemma-7B menjadi 5%. Pola serupa muncul pada penolakan tugas mustahil, yaitu 86% berbanding 23% dan 0%. Korelasi terbalik yang konsisten pada kedua model dasar tersebut mengindikasikan bahwa multiplikasi jumlah parameter secara mandiri belum memadai untuk mengeksplorasi penalaran spasial yang terikat realitas fisik (grounded). Kelayakan arsitektur dikonfirmasi melalui dekomposisi latensi: inferensi server menyumbang 88,35% sampai 99,92% dari total waktu respons, sementara transfer jaringan pada klien hanya 10 sampai 51 ms, atau 1,16% sampai 2,10%. Karena itu, sistem bersifat tergantung pada komputasi, sehingga arsitektur terdistribusi merupakan konfigurasi yang optimal dan bukan sekadar kompromi. Kelayakan operasional sisi server diperkuat oleh profil daya dan termal selama 800 kasus uji berkelanjutan: konsumsi daya GPU bertahan stabil pada 180 sampai 200 W tanpa penurunan transfer data, dengan suhu operasional pada 60 sampai 65 ◦C, jauh di bawah ambang penurunan frekuensi pabrikan sebesar 83 ◦C, sehingga waktu tanggap ke lapisan edge tetap konsisten sepanjang beban penuh. Pada Level 1, lokalisasi Ultrawideband mencapai galat laterasi 98,88 mm pada kondisi Line of Sight dan 279,94 mm pada Non-Line of Sight, dengan reduksi galat pembacaan jarak tempuh dari 14,36% menjadi 1,096%, sementara persepsi gestural mencapai akurasi 93,5% dengan model berukuran 7 KB. Pada Level 3, navigasi lintas-zona mencapai keberhasilan 100%. menekan galat titik henti robot terhadap titik tujuan dari rentang 24,52 sampai 33,42 cm menjadi 15,47±6,39 cm pada 20 percobaan, yaitu sekitar separuhnya, sekaligus memangkas waktu fase pendekatan dari rentang 29 sampai 49 detik menjadi 14,7 sampai 16,6 detik. Manipulasi objek berhasil pada 45 dari 60 percobaan, yaitu 75,0%. Seluruh kegagalan bersumber pada keterbatasan perangkat keras, terdiri atas 11 kegagalan mekanis pada penjepit dan jangkauan lengan serta 4 kegagalan akibat degradasi data sensor kedalaman, dan tidak satu pun bersumber pada lapisan penalaran. Klaim arsitektural yang diajukan bersifat umum terhadap jenis konteks, dan telah terverifikasi pada dua modalitas yang berbeda, yaitu persepsi visual objek dan lokalisasi spasial berbasis SLAM. Pembuktian langsung pada modalitas Ultrawideband, melalui integrasi end-to-end antara jaringan sensor dan lapisan deliberatif belum dilaksanakan secara penuh sehingga masih diperlukan pengujian lanjutan.
==================================================================================================================================
The growing elderly population demands service robots capable of natural, language-based collaboration. However, the computational load of Large Language Models (LLMs) exceeds the capacity of onboard edge devices, forcing the reasoning layer to be physically separated onto a server. This separation gives rise to the derivative problem at the core of this dissertation: a reasoning layer distanced from the robot loses access to the state of the physical world, while instructions from elderly users are typically incomplete and heavily context-dependent. This dissertation advances the thesis that ambient intelligence is not a supplementary feature but an architectural precondition that enables the distribution of the deliberative layer without loss of grounding. The proposed architecture comprises three levels: Level 1 (ambient intelligence at the edge, encompassing Ultrawideband localization and gestural perception), which supplies a context vector; Level 2 (cognitive reasoning on a server, using a QLoRAfine- tuned Gemma-2B with human-in-the-loop validation), which maps utterance and context into a structured plan; and Level 3 (physical execution back at the edge, encompassing navigation and manipulation), which translates the symbolic plan into motor actions. Validation follows a staged-combination strategy to localize error sources at each level. The most decisive finding emerges from a cross-model comparison conducted under an identical schema-constrained prompt. On direct commands that require no spatial grounding, both substantially larger baselines lead, with GPT-3.5 reaching 97% and Gemma-7B 94%, against 89% for the fine-tuned Gemma-2B. That advantage reverses on spatial reasoning, where Gemma-2B sustains 87% while GPT-3.5 falls to 15% and Gemma-7B to 5%. The same pattern recurs on impossible-task rejection, at 86% against 23% and 0%. Because the reversal holds consistently across both baselines, added parameter capacity by itself does not yield grounded spatial reasoning. Architectural feasibility is confirmed through latency decomposition: server inference accounts for 88.35% to 99.92% of total response time, while network transfer on a Raspberry Pi 5 client contributes only 10 to 51 ms, or 1.16% to 2.10%. The system is thus computation-bound, making the distributed architecture an optimal configuration rather than a mere compromise. Serverside operational feasibility is further supported by power and thermal profiling over 800 sustained test cases: GPU power draw remains stable at 180 to 200 W with no I/O blocking, and operating temperature holds at 60 to 65 ◦C, far below the 83 ◦C throttling threshold, so the response time supplied to the edge layer stays consistent under full load. At Level 1, Ultrawideband localization attains a lateration error of 98.88 mm under Line of Sight and 279.94 mm under Non-Line of Sight, with the reading error over the travelled distance reduced from 14.36% to 1.096%, while gestural perception attains 93.5% accuracy with a 7 KB model. At Level 3, cross-zone navigation succeeds in 100% of trials. Replacing the final approach phase from open-loop kinematic estimation with closed-loop Image-Based Visual Servoing reduces the stopping-point error relative to the target from a range of 24.52 to 33.42 cm down to 15.47 ± 6.39 cm over 20 trials, roughly half the earlier figure, while cutting the approach phase from 29–49 seconds to 14.7–16.6 seconds and removing the bounding-box redrawing previously required of the operator whenever the object scale changed. Object manipulation succeeds in 45 of 60 trials, or 75.0%. All failures trace to hardware limitations, comprising 11 mechanical failures of the gripper and arm reach and 4 failures caused by depth-sensor degradation, with none originating in the reasoning layer. The architectural claims advanced here are general with respect to the type of context, and have been verified across two fundamentally different modalities, namely object-level visual perception and SLAM-based spatial localization. Direct demonstration on the Ultrawideband modality, through end-to-end integration between the anchor network and the deliberative layer has not been fully implemented and requires further testing.

Item Type: Thesis (Doctoral)
Uncontrolled Keywords: Arsitektur Terdistribusi, Kolaborasi Manusia-Robot, Kecerdasan Ambien, Model Bahasa Besar, Symbol Grounding, Komputasi Edge. Distributed Architecture, Human-Robot Collaboration, Ambient Intelligence, Large Language Models, Symbol Grounding, Edge Computing.
Subjects: T Technology > T Technology (General) > T59.7 Human-machine systems.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Electrical Engineering > 20001-(S3) PhD Thesis
Depositing User: Muhtadin Muhtadin
Date Deposited: 07 Aug 2026 02:26
Last Modified: 07 Aug 2026 02:26
URI: http://repository.its.ac.id/id/eprint/144194

Actions (login required)

View Item View Item