Beyond Taste: Penilaian Estetika Fotografi Makanan Menggunakan Prompting Multimodal

Farzana, Aria Nalini (2026) Beyond Taste: Penilaian Estetika Fotografi Makanan Menggunakan Prompting Multimodal. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221016-Undergraduate_Thesis.pdf] Text
5025221016-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (8MB) | Request a copy

Abstract

Penilaian estetika fotografi makanan umumnya hanya berfokus pada pemanfaatan fitur visual, padahal penggunaan informasi tekstual berpotensi untuk merepresentasikan deskripsi visual dengan bahasa natural. Penelitian ini bertujuan untuk mengimplementasikan dan membandingkan teknik prompting kontrastif dan terstruktur dalam skenario general dan detail dalam pembuatan AI-Generated caption serta mengevaluasi pengaruhnya terhadap representasi tekstual pada kinerja sistem penilaian estetika multimodal fotografi makanan. Penelitian ini menggunakan Gourmet Photography Dataset (GPD) berisi pasangan gambar makanan dengan label estetika biner High Aesthetic dan Low Aesthetic. Large Language and Vision Assistant (LLaVA) digunakan untuk generasi caption dengan empat skenario prompting: kontrastif general, kontrastif detail, terstruktur general, dan terstruktur detail. Contrastive Language-Image Pre-Training (CLIP) digunakan untuk mengekstraksi fitur visual dan tekstual yang digabungkan menggunakan Fusion Multilayer Perceptron untuk klasifikasi estetika biner. Hasil penelitian mengungkapkan bahwa teknik prompting memengaruhi panjang, struktur, dan tingkat detail caption, namun tidak memberikan perbedaan kinerja yang signifikan. Model multimodal terbaik memperoleh Accuracy sebesar 91,56%, F1-Score 92,35%, dan Balanced Accuracy sebesar 91,36%, sedikit lebih tinggi dibanding model visual only yang memperoleh Accuracy sebesar 91,31%, F1-Score sebesar 92,21%, dan Balanced Accuracy sebesar 90,99%. Temuan ini menunjukkan bahwa informasi visual tetap menjadi sumber informasi utama, sedangkan informasi tekstual berperan sebagai informasi tambahan yang dapat memberikan peningkatan kinerja model klasifikasi secara konsisten meskipun relatif terbatas.
===============================================================================================================================
Aesthetic evaluation of food photography generally focuses only on the use of visual features, even though textual information has the potential to represent visual descriptions in natural language. This study aims to implement and compare contrastive and structured prompting techniques in general and detailed scenarios for generating AI-generated captions, as well as to evaluate their impact on textual representations in the performance of a multimodal aesthetic evaluation system for food photography. This study uses the Gourmet Photography Dataset (GPD), which contains pairs of food images labeled with the binary aesthetic categories “High Aesthetic” and “Low Aesthetic.” The Large Language and Vision Assistant (LLaVA) is used for caption generation across four prompting scenarios: contrastive-general, contrastive-detail, structured-general, and structured-detail. Contrastive Language-Image Pre-Training (CLIP) was used to extract visual and textual features, which were then combined using a Fusion Multilayer Perceptron for binary aesthetic classification. The results revealed that prompting techniques influenced the length, structure, and level of detail in the captions but did not yield significant differences in performance. The best multimodal model achieved an accuracy of 91.56%, an F1-score of 92.35%, and a balanced accuracy of 91.36%, which is slightly higher than the visual-only model, which achieved an accuracy of 91.31%, an F1-score of 92.21%, and a balanced accuracy of 90.99%. These findings indicate that visual information remains the primary source of information, while textual information serves as supplementary information that can consistently improve the performance of classification models, albeit to a relatively limited extent.

Item Type: Thesis (Other)
Uncontrolled Keywords: CLIP, Fotografi Makanan, Image Aesthetic Assessment, LLaVA, Multimodal Visual-Teks, Prompting, CLIP, Food Photography, Image Aesthetic Assessment, LLaVA, Prompting, Visual-Text Multimodal.
Subjects: Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines.
T Technology > T Technology (General) > T58.6 Management information systems
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Aria Nalini Farzana
Date Deposited: 28 Jul 2026 01:00
Last Modified: 28 Jul 2026 01:00
URI: http://repository.its.ac.id/id/eprint/137934

Actions (login required)

View Item View Item