fleiss-kappa-af472c10·2 events·first seen Aliases: Fleiss' Kappa, Fleiss kappa
Researchers present a suite of three small language models (146M to 3B parameters) built on a hyperbolic geometric substrate, targeting trustworthy companion AI. A 146M behavioral auditor achieves 90.7% binary-compliance accuracy and outperforms a frontier zero-shot judge (AUROC 0.804 vs 0.721) in detecting sycophancy, dependence-fostering, and confabulated memories across unseen generator families. A creative frame-seeder wins 100% of 311 pairwise comparisons over prompting baselines, and a memory OS implements exponential decay-based 'designed forgetting.' The work proposes a small-model alternative to frontier-scale approaches for companion AI safety and personalization.
This paper introduces a large, consensus-labeled benchmark of 6,675 prompts drawn from eight existing corpora (ASTRA, CySecBench, AdvBench, JailbreakBench, MalwareBench, RedCode, RMCBench, Scam2Prompt) to evaluate whether coding-specialized LLMs refuse malicious requests. A key contribution is the distinction between requests for executable malicious code (4,748 prompts) versus harmful security knowledge (1,923 prompts), arguing that coding models should face a stricter refusal standard given their outputs can be directly weaponized. A five-judge consensus protocol achieves Fleiss' kappa of 0.767, providing a reliability-quantified substrate for cross-corpus compliance measurement that the field has previously lacked.