eprintid: 78779 rev_number: 10 eprint_status: archive userid: 12460 dir: disk0/00/07/87/79 datestamp: 2026-10-07 02:05:38 lastmod: 2026-10-07 02:05:38 status_changed: 2026-10-07 02:05:38 type: thesis metadata_visibility: show contact_email: muh.khabib@uin-suka.ac.id creators_name: Subhan Ghufron, NIM.: 24206051015 title: OPTIMALISASI ARSITEKTUR RETRIEVAL-AUGMENTED GENERATION MELALUI FINE-TUNING MODEL EMBEDDING DAN LARGE LANGUAGE MODEL UNTUK ANALISIS NAHWU-SHARAF KITAB KUNING ispublished: pub subjects: 006.3 divisions: S2_inf full_text_status: restricted keywords: Retrieval-Augmented Generation, Nahwu-Sharaf, Kitab Kuning, Fine-tuning, Model Embedding, Large Language Model, BGE-M3, E5-Large, Llama3-8B, Mistral-7B note: Prof. Dr. Ir. Shofwatul 'Uyun,S.T.,M.Kom., IPM., ASEAN Eng. abstract: Arabic grammatical analysis (Nahwu-Sharaf) of classical Islamic texts (kitab kuning) is a cornerstone of the pesantren education tradition in Indonesia. This process still relies on manual analysis requiring deep expertise and extensive time. This study aims to evaluate the effectiveness of fine-tuning embedding models and Large Language Models (LLMs) within a Retrieval-Augmented Generation (RAG) architecture for automated Nahwu-Sharaf analysis of kitab kuning. Twelve configurations were tested comparatively, involving two embedding models (BGE-M3 and E5-Large) and two LLMs (Llama3-8B and Mistral-7B), each in baseline and fine-tuned conditions using QLoRA on the NahwuAI dataset. Evaluation was conducted using automatic metrics (Recall@K, BLEU, ROUGE-L, F1, Cosine Similarity) and human evaluation by three Arabic grammar experts. Results showed that fine-tuned BGE-M3 outperformed as the embedding model with a Recall@1 of 0.8109 (a 37.7% increase). Fine-tuned Mistral-7B was selected as the best LLM due to its stronger generalization capability on kitab kuning texts compared to Llama3-8B, which tended to overfit to training data. The RAG + fine-tuned Mistral configuration achieved the highest F1 (0.1176) and the highest expert score (3.80 out of 5). RAG was demonstrated to provide the greatest added value when paired with a model already strengthened through fine-tuning. date: 2026-08-10 date_type: published pages: 117 institution: UIN SUNAN KALIJAGA YOGYAKARTA department: FAKULTAS SAINS DAN TEKNOLOGI thesis_type: masters thesis_name: other citation: Subhan Ghufron, NIM.: 24206051015 (2026) OPTIMALISASI ARSITEKTUR RETRIEVAL-AUGMENTED GENERATION MELALUI FINE-TUNING MODEL EMBEDDING DAN LARGE LANGUAGE MODEL UNTUK ANALISIS NAHWU-SHARAF KITAB KUNING. Masters thesis, UIN SUNAN KALIJAGA YOGYAKARTA. document_url: https://digilib.uin-suka.ac.id/id/eprint/78779/1/24206051015_BAB-I_IV-atau-V_DAFTAR-PUSTAKA.pdf document_url: https://digilib.uin-suka.ac.id/id/eprint/78779/2/24206051015_BAB-II_sampai_SEBELUM-BAB-TERAKHIR.pdf