retrieval-augmented-fine-tuning-e49ab2eb·1 events·first seen Aliases: Retrieval-Augmented Fine-Tuning
Researchers benchmark eight fine-tuning and retrieval configurations of Gemma 4 31B-IT against the U.S. NRC Reactor Operator licensing examination across 14 Generic Fundamentals Examinations from 2015–2021. The best configuration—supervised fine-tuning with Gemini-distilled chain-of-thought rationales combined with fixed-size chunking RAG—passed 8 of 14 exams and reached 79.7% aggregate accuracy, near the 80% human passing threshold. Key findings include that no configuration without fine-tuning passed any exam, that preferred chunking strategy reverses depending on training state, and that RAFT underperforms standard SFT in matched retrieval environments.