A preprint from arXiv proposes and validates an architectural approach combining locally deployed LLMs with Retrieval-Augmented Generation (RAG) for regulatory compliance and legal document analysis, specifically targeting on-premises deployment without high-end GPU hardware. The implementation uses Ollama and LM Studio as execution environments with Polish-language models Bielik and PLLuM running on consumer-class hardware. Results show RAG integration improves factual consistency, domain specificity, and normative precision while enabling auditability and dynamic knowledge updates without retraining.
Hugging Face published a case study describing how Digital Green used an LLM-as-a-Judge approach to evaluate and improve a retrieval-augmented generation (RAG) application. The post covers the methodology for using LLMs to score and validate RAG outputs, providing a practical deployment pattern for quality assurance in production AI systems. It serves as a concrete example of enterprise-grade evaluation pipelines built on top of RAG architectures.
Mistral AI published a technical guide on evaluating Retrieval-Augmented Generation (RAG) systems using the 'LLM as a Judge' paradigm combined with their structured outputs API feature. The approach implements the RAG Triad framework—context relevance, groundedness, and answer relevance—using Pydantic schemas to enforce machine-readable evaluation outputs. Mistral models serve as both the generator and judge components, enabling scalable automated evaluation without human annotators.
A curated GitHub repository collecting over 100 deployable AI agent and RAG (Retrieval-Augmented Generation) applications built with LLMs. The collection is designed for practical use — clone, customize, and ship. With 110,915 total stars and 202 added today, it reflects strong community interest in applied LLM tooling.
A new arXiv paper presents the first systematic study of using reinforcement learning to teach LLMs to adapt query formulation strategies to different retrieval backends. The authors find that different retrievers have surprisingly distinct optimal query styles (e.g., descriptive vs. question-like), making cross-retriever strategy transfer ineffective. They introduce a branching-based rollout technique to stabilize training over multi-step retrieval trajectories and show gains from retriever-specific human guidance and model scaling.
A preprint from arXiv proposes a theoretical framework for embedding cognitive architecture natively into LLMs rather than simulating it via prompt engineering and context management. The framework introduces three mechanisms: Structural Tension (an endogenous loss function from information-manifold conflict), an Offline Recurrent Loop (sandboxed self-processing without external input), and Inference-time Plasticity (topology reconfiguration without weight modification). The authors argue these mechanisms could produce heterogeneous model instances with distinct topological structures through path-dependent evolution, while remaining within governance rails. The paper is primarily theoretical, offering operational definitions, reconfiguration operators, and falsification criteria rather than empirical results.
A new arXiv preprint proposes a three-agent framework for sanitizing retrieved content in RAG pipelines by performing privacy extraction, semantic analysis, and reconstruction as an offline preprocessing step. Evaluated on ChatDoctor and Wiki-PII datasets across six LLMs, the approach reduces targeted information exposure in LLaMA-3-8B from 144 baseline instances to 1, while maintaining contextual fidelity (BLEU-1 of 0.122 vs. SAGE's 0.117). The framework introduces no additional online inference latency since rewriting is done offline. Source code is publicly released.
A case study on the Danish National Encyclopedia's RAG system evaluates five retrieval workflows across 20,000 query-workflow pairs, revealing a 'Coverage Illusion' where synthetic queries overestimate the need for LLM augmentation (90%+) versus real production traffic (27.8%). Pre-retrieval routing cannot detect this gap because augmentation necessity is only revealed after index search. A post-retrieval cascade running workflows cheapest-first and escalating to LLM augmentation only on empty results improves quality by +0.140 Composite Overall points over Always-HyDE, reduces latency by 31.8%, and eliminates LLM augmentation for 72.2% of real queries. The work highlights a structural mismatch between synthetic and real query distributions that affects RAG system design assumptions.
Researchers introduce PAT (Pragmatic Auto-Translator), a RAG-based system that moves LLM translation beyond sentence-by-sentence processing by pairing user-configured specifications with paragraph-, section-, and document-level examples from a comparable corpus of U.S. English and Latin American Spanish texts. The system targets draft translation for professional verification, aiming to reformulate discourse organization, rhetorical style, and pragmatic norms for the target language context. Evaluation using a customized MQM typology across six automatic translations of generative AI essays found that limited prompting produced no meaningful reformulation, while specifications and corpus-informed approaches showed partial but inconsistent improvement. The work demonstrates LLMs can be nudged toward document-level reformulation but highlights remaining gaps in effectiveness.