dream-a254a3ee·2 events·first seen Aliases: Dream
DREAM is a new method for training dense retrieval embedding models using the autoregressive next-token prediction objective of a frozen LLM, bypassing the need for labeled positive/negative document pairs required by contrastive training. The approach injects retriever-generated query-document similarity scores into selected attention heads of the LLM, allowing prediction loss gradients to flow back to the retriever. Evaluated on BEIR and RTEB benchmarks with 0.5B–3B parameter backbones, DREAM consistently outperforms contrastive baselines across model scales.
A new arXiv paper investigates whether locate-then-edit knowledge editing methods, developed for autoregressive models, transfer to masked diffusion language models (MDMs) such as LLaDA and Dream. The authors find that causal tracing identifies the same early-to-mid-layer MLP location in both paradigms, but MDMs degrade systematically on multi-token edits due to partially unmasked intermediate states that the edit was never optimized for. A correction targeting these intermediate states substantially restores multi-token editing performance. The work is the first systematic comparison of knowledge editing across autoregressive and diffusion-based language model paradigms.