Kerangka Kerja Cognitive Chain-of-thought (Cog-CoT) Berbasis Dekomposisi Kognitif Menggunakan Taksonomi Bloom Untuk Generasi Jawaban Edukatif

Muzli, Alfarabi (2026) Kerangka Kerja Cognitive Chain-of-thought (Cog-CoT) Berbasis Dekomposisi Kognitif Menggunakan Taksonomi Bloom Untuk Generasi Jawaban Edukatif. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 6025241061-Master_Thesis.pdf] Text
6025241061-Master_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (4MB) | Request a copy

Abstract

Large Language Model (LLM) menghadapi tantangan Cognitive Misalignment dalam konteks pendidikan formal, yaitu kondisi ketika jawaban yang dihasilkan secara linguistik meyakinkan namun secara pedagogis tidak sesuai dengan kedalaman kognitif yang dituntut oleh pertanyaan berdasarkan Taksonomi Bloom. Selain itu, arsitektur Retrieval-Augmented Generation (RAG) yang ada hanya memverifikasi output akhir, sehingga model dapat menyimpang dari konteks faktual selama proses reasoning berlangsung. Penelitian ini mengusulkan Cognitive Chain-of-Thought (Cog-CoT), sebuah kerangka kerja empat modul terintegrasi yang mengoperasionalkan Taksonomi Bloom sebagai scaffold kognitif formal dalam generasi jawaban edukatif, terdiri atas (1) Cognitive Classification, (2) Cognitive Decomposition, (3) Cog-CoT Reasoning dengan mekanisme self-verification berbasis cosine similarity dan retry, serta (4) Bottom-Up Answer Aggregation. Eksperimen dilakukan pada domain Ilmu Pengetahuan Sosial (IPS) jenjang SMP-SMA dalam bahasa Indonesia, menggunakan tiga model LLM open-weight yaitu Gemma-3-4B-IT, LLaMa-3.1-8B-Instruct, dan Qwen2.5-7B-Instruct. Cognitive Classifier mencapai akurasi 99,15% menggunakan LinearSVC. Cog-CoT secara konsisten unggul dibandingkan Zero-Shot dan Standard CoT pada seluruh model, dengan GT cosine similarity tertinggi pada Gemma sebesar 0,8066, LLaMa sebesar 0,8083, dan Qwen sebesar 0,8124, atau peningkatan 6,90–7,36% terhadap Zero-Shot, yang dikonfirmasi signifikan secara statistik melalui uji Wilcoxon signed-rank pada 55 dari 60 pengujian (91,7%, p < 0,05). Penelitian ini juga menemukan fenomena penurunan performa CoT, yaitu Standard CoT justru menghasilkan performa di bawah Zero-Shot pada seluruh model akibat semantic drift tanpa panduan kognitif, sebuah fenomena yang juga terkonfirmasi pada dataset benchmark eksternal SQuAD 2.0 dan ASQA. Studi ablasi menunjukkan bahwa modul RAG dan self-verification saling bergantung secara fungsional, sementara dekomposisi kognitif memberikan manfaat terbesar pada level penalaran tingkat tinggi (C4–C6). Keunggulan Cog-CoT pada benchmark eksternal membuktikan sifat model-agnostic dan generalizability-nya lintas domain dan bahasa.
=======================================================================================================================================
Large Language Models (LLMs) face a fundamental challenge in formal educational contexts known as Cognitive Misalignment, a condition in which generated responses are linguistically convincing yet pedagogically misaligned with the cognitive depth demanded by the question according to Bloom's Taxonomy. Moreover, existing Retrieval-Augmented Generation (RAG) architectures verify only the final output, leaving models free to drift from factual context during intermediate reasoning steps. This study proposes Cognitive Chain-of-Thought (Cog-CoT), a four-module integrated framework that operationalizes Bloom's Taxonomy as a formal cognitive scaffold for educational answer generation, comprising (1) Cognitive Classification; (2) Hierarchical Cognitive Decomposition; (3) Cog-CoT Reasoning, with a cosine-similarity-based self-verification and retry mechanism; and (4) Bottom-Up Answer Aggregation. Experiments were conducted in the Indonesian Social Studies (IPS) domain for junior and senior high school levels, using three open-weight LLMs: Gemma-3-4B-IT, LLaMa-3.1-8B-Instruct, and Qwen2.5-7B-Instruct. The Cognitive Classifier achieved 99.15% accuracy using LinearSVC. Cog-CoT consistently outperformed Zero-Shot and Standard CoT across all models, achieving the highest GT Cosine similarity on Gemma (0.8066), LLaMa (0.8083), and Qwen (0.8124), an improvement of 6.90-7.36% over Zero-Shot, confirmed as statistically significant by Wilcoxon signed-rank tests in 55 of 60 comparisons (91.7%, p < 0.05). The study further identifies a CoT degradation phenomenon in which Standard CoT consistently underperforms Zero-Shot across all models due to semantic drift, a pattern also confirmed on the external benchmark datasets SQuAD 2.0 and ASQA. Ablation studies show that the RAG and self-verification modules are functionally interdependent, while cognitive decomposition provides the greatest benefit at higher-order reasoning levels (C4-C6). The superiority of Cog-CoT on these external benchmarks further demonstrates its model-agnostic nature and generalizability across domains and languages.

Item Type: Thesis (Masters)
Uncontrolled Keywords: Chain-of-Thought prompting, Cognitive Decomposition, Cognitive Misalignment, generasi jawaban edukatif, Large Language Model, RAG, self-verification, Taksonomi Bloom, Bloom’s Taxonomy, educational answer generation, Large Language Model,
Subjects: Q Science > QA Mathematics > QA336 Artificial Intelligence
Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: Alfarabi Muzli
Date Deposited: 20 Jul 2026 01:59
Last Modified: 20 Jul 2026 01:59
URI: http://repository.its.ac.id/id/eprint/135464

Actions (login required)

View Item View Item