technique

SFT

techniqueactiveprovisionalsft-cfebfef1·1 events·first seen 37h ago

Aliases: SFT

Co-occurring entities

GRPO DPO Can LLMs Reliably Self-Report Adversarial Prefills, and How?

More like this (12)

Target-SFT Target-SFT FedTSV SIFT FTX MemFT ChunkFT SiT DeltaFS SGSD SD-Turbo SETA

Recent events (1)

6arXiv · cs.CL·37h ago·source ↗

LLMs fail to reliably self-report adversarial prefill attacks, study finds

A new arXiv paper evaluates whether LLMs can recognize that their own prior responses were elicited by adversarial prefill attacks, testing ten open-weight models (3B–70B) across four safety benchmarks. Models claim intent on prefilled responses only 27.3% of the time on average, and introspective signal is largely mediated by refusal-related reasoning. Three LoRA fine-tuning methods (SFT, GRPO, DPO) improve the intention-probe gap but counterintuitively raise attack success rates on most models, suggesting partial and fragile mitigation. The findings raise concerns about the reliability of LLM self-reports in safety-critical contexts.

Evaluation and Benchmarking AI Safety Research GRPO DPO Can LLMs Reliably Self-Report Adversarial Prefills, and How?+1 more