Pratama, Fanza Khairan (2026) Image Captioning Makanan dan Nutrisi Berbasis Prediksi CNN dan Transformer pada Dataset Nutrition5k. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221305-Undergraduate_Thesis.pdf Restricted to Repository staff only Download (9MB) | Request a copy |
Abstract
Pemantauan asupan nutrisi berbasis citra makanan merupakan permasalahan penting dalam konteks kesehatan preventif dan manajemen pola makan, karena manusia sering mengalami kesulitan dalam mengestimasi kandungan gizi hanya dari pengamatan visual. Perkembangan deep learning telah mendorong munculnya pendekatan visual nutrition understanding yang mampu memprediksi nilai nutrisi dari citra makanan secara otomatis. Namun, sebagian besar penelitian masih berfokus pada keluaran numerik sehingga informasi yang dihasilkan kurang komunikatif bagi pengguna awam. Di sisi lain, pendekatan image captioning generatif mampu menghasilkan deskripsi bahasa alami tetapi sering mengalami hallucination faktual, khususnya pada domain nutrisi yang menuntut konsistensi terhadap nilai yang terukur secara fisik. Penelitian ini mengembangkan sistem nutritional image captioning berbasis fakta yang mengintegrasikan prediksi nutrisi CNN dengan pembangkitan deskripsi bahasa alami menggunakan arsitektur Transformer. Sistem dirancang secara bertahap dengan memisahkan secara eksplisit proses estimasi nutrisi dari proses pembangkitan teks untuk meminimalkan risiko hallucination faktual. Dataset Nutrition5k digunakan sebagai basis eksperimen, menyediakan citra overhead RGB makanan beserta anotasi nutrisi berbasis pengukuran fisik yang mencakup kalori, protein, lemak, karbohidrat, dan massa. Nilai kelima atribut nutrisi diprediksi secara simultan menggunakan CNN berbasis EfficientNetV2B0, kemudian dilinearisasi menjadi fact string sebagai kondisi bagi sistem pembangkitan teks berbasis ViT-B/16 dan T5-small melalui mekanisme encoder-side concatenation. Enam model dibandingkan dalam eksperimen, meliputi dua baseline, dua model fact-grounded, dan dua model konvensional tanpa fact grounding. Hasil evaluasi menunjukkan bahwa model CNN berhasil memprediksi atribut nutrisi dengan R² berkisar antara 0,696 hingga 0,891. Model fact-grounded utama (ViT+T5 Freeze) mencapai BLEU-4 sebesar 0,4062 dan CIDEr sebesar 2,3827. Pada dimensi konsistensi faktual, model fact-grounded mencapai Nutritional Accuracy (NutAcc) rata-rata sebesar 96,37% berbanding 16,57% pada model konvensional, dengan selisih 79,80 poin persentase. MAE kalori terhadap ground truth fisik juga turun dari 127,17 kcal menjadi 52,25 kcal, membuktikan bahwa mekanisme fact grounding secara struktural mampu menekan hallucination numerik pada domain nutrisi makanan.
=======================================================================================================================================
Monitoring dietary intake from food images is an important problem in preventive healthcare and dietary management, as humans often struggle to accurately estimate nutritional content based solely on visual perception. Recent advances in deep learning have enabled visual nutrition understanding approaches that automatically predict nutritional values from food images. However, most existing studies focus primarily on numerical outputs, limiting their interpretability for general users. Meanwhile, generative image captioning approaches can produce natural language descriptions but frequently suffer from factual hallucination, particularly in nutrition-related domains that require strong consistency with physically measured values. This research develops a fact-grounded nutritional image captioning system that integrates CNN-based nutrient prediction with natural language description generation using a Transformer architecture. The system is designed in a staged manner that explicitly separates nutrient estimation from text generation to minimize the risk of factual hallucination. The Nutrition5k dataset is used as the experimental basis, providing overhead RGB food images with nutrition annotations obtained through direct physical measurement, covering calories, protein, fat, carbohydrates, and mass. All five nutritional attributes are predicted simultaneously using an EfficientNetV2B0-based CNN, then linearized into a fact string that conditions the text generation component based on ViT-B/16 and T5-small through an encoder-side concatenation mechanism. Six models are compared in the experiments, comprising two baselines, two fact-grounded models, and two conventional models without fact grounding. Experimental results show that the CNN model achieves R² values ranging from 0.696 to 0.891 across all predicted nutritional attributes. The primary fact-grounded model (ViT+T5 Freeze) achieves a BLEU-4 score of 0.4062 and a CIDEr score of 2.3827. On the factual consistency dimension, the fact-grounded model achieves an average Nutritional Accuracy (NutAcc) of 96.37% compared to 16.57% for the conventional model, a difference of 79.80 percentage points. The calorie MAE against physical ground truth also decreases from 127.17 kcal to 52.25 kcal, demonstrating that the fact grounding mechanism structurally suppresses numerical hallucination in the food nutrition domain.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Nutritional Image Captioning, Visual Nutrition Understanding, Convolutional Neural Network, Vision Transformer, T5, Fact Grounding, Hallucination |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science) T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing. |
| Divisions: | Faculty of Information Technology > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Fanza Khairan Pratama |
| Date Deposited: | 28 Jul 2026 04:16 |
| Last Modified: | 28 Jul 2026 04:16 |
| URI: | http://repository.its.ac.id/id/eprint/138582 |
Actions (login required)
![]() |
View Item |
