Augmentasi Data Citra Postur Duduk Menggunakan Pipeline Generatif Stable Diffusion XL dan Motion In-betweening untuk Evaluasi Ergonomi

Jonathan, Billy (2026) Augmentasi Data Citra Postur Duduk Menggunakan Pipeline Generatif Stable Diffusion XL dan Motion In-betweening untuk Evaluasi Ergonomi. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221170-Undergraduate_Thesis.pdf] Text
5025221170-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (8MB) | Request a copy

Abstract

Asesmen risiko postur ergonomi secara otomatis membutuhkan dataset video berskala besar yang sulit diperoleh secara manual di lapangan. Penelitian ini mengembangkan pipeline augmentasi data sintetis berbasis Latent Diffusion Model (LDM) untuk memproduksi video postur kerja untuk mengatasi kendala class imbalance pada dataset pelatihan, terutama untuk kategori variasi postur duduk ekstrem yang berisiko tinggi. Proses diawali dengan ekstraksi pose menggunakan DWPose, yang dilanjutkan dengan manipulasi dan sintesis gerak berbasis kecerdasan buatan menggunakan model Transformer untuk motion in-betweening yang dibatasi oleh standar rentang gerak anatomis NASA. Untuk memitigasi inkonsistensi identitas dan fluktuasi latar belakang pada video, dilakukan studi ablasi pada arsitektur Stable Diffusion XL dengan mengintegrasikan ControlNet, Image Prompt Adapter (IP-Adapter), Low-Rank Adaptation (LoRA), serta teknik Surgical Compositing. Hasil evaluasi menunjukkan bahwa integrasi penuh IP-Adapter Plus dan LoRA memberikan performa paling optimal dengan skor Percentage of Correct Keypoints (PCK) sebesar 68,99%, Mean Absolute Error (MAE) sebesar 12,62°, temporal Structural Similarity Index (t-SSIM) dengan skor 0,963, dan JEDi score sebesar 2,687. Pengujian pada model klasifikasi regresi Bi-LSTM menunjukkan adanya penurunan tipis pada performa baseline (Validasi R2 turun dari 0,876 ke 0,871) akibat jitter dari citra sintetis yang dibaca sebagai noise oleh Pose Estimator. Meskipun masih terdapat kendala concept bleeding pada unseen data dan anomali hilangnya objek, model yang diusulkan efektif menghasilkan video postur yang presisi secara anatomi dan konsisten secara visual untuk kebutuhan klasifikasi biomekanika.
===========================================================================================================================================
Automated ergonomic posture risk assessment requires large-scale video datasets that are difficult to obtain manually in the field. This study develops a synthetic data augmentation pipeline based on a Latent Diffusion Model (LDM) to generate work posture videos to address the class imbalance issues in training datasets, particularly for high-risk, extreme sitting posture variations. The process begins with pose extraction using DWPose, followed by AI-based motion manipulation and synthesis using a Transformer model for motion in-betweening, constrained by NASA's anatomical range of motion standards. To mitigate identity inconsistencies and background fluctuations in the generated videos, an ablation study was conducted on the Stable Diffusion XL architecture by integrating ControlNet, Image Prompt Adapter (IP-Adapter), Low-Rank Adaptation (LoRA), and a Surgical Compositing technique. Evaluation results show that the full integration of IP-Adapter Plus and LoRA provides the most optimal performance, achieving a Percentage of Correct Keypoints (PCK) score of 68.99%, a Mean Absolute Error (MAE) of 12.62°, a temporal Structural Similarity Index (t-SSIM) score of 0.963, and a JEDi score of 2.687. Testing on the Bi-LSTM regression classifier model indicates a marginal decline in baseline performance (validation R2 decreased from 0.876 to 0.871) due to jitter in the synthetic images being interpreted as noise by the pose estimator. Although limitations such as concept bleeding on unseen data and object disappearance anomalies persist, the proposed model is effective in generating anatomically precise and visually consistent posture videos for biomechanics classification needs.

Item Type: Thesis (Other)
Uncontrolled Keywords: Augmentasi Data, IP-Adapter, Latent Diffusion Model, LoRA, Motion In-betweening, Sintesis Video, Data Augmentation, IP-Adapter, Latent Diffusion Model, LoRA, Motion In-betweening, Video Synthesis
Subjects: Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Q Science > QP Physiology > QP34.5 Human engineering
T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Billy Jonathan
Date Deposited: 28 Jul 2026 02:15
Last Modified: 28 Jul 2026 02:15
URI: http://repository.its.ac.id/id/eprint/137993

Actions (login required)

View Item View Item