ragtruth-7c6fc9a2·1 events·first seen Aliases: RAGTruth
Researchers introduce D-Score, a hallucination detection method based on the spectral geometry of LLM hidden activations, computed in a single forward pass without external verifiers, retrieval, or multiple generations. The method counts singular directions of the hidden activation matrix whose singular values remain close to the leading one, operating on the intuition that hallucinated text causes representations to spread across additional singular directions. Evaluation on FAVA-Annotation and RAGTruth benchmarks shows strong detection performance. The approach is lightweight and model-intrinsic, making it practically attractive for deployment.