Researchers introduce 'appearance pointers,' compact tokens that enable precise regional control in Diffusion Transformers (DiTs) by aligning text or image inputs with user-specified spatial masks. A region correspondence network produces these tokens, refined via spatial aggregation, allowing multi-region guidance without retraining the base model or significantly increasing token load. The method claims to be the first modality-agnostic interface for localized multimodal control in DiTs, matching or surpassing modality-specific state-of-the-art methods across evaluated metrics.
TunerDiT is a training-free method for steering video diffusion transformers (DiTs) to generate long-horizon videos containing multiple sequential events. The approach identifies intrinsic turning points in the DiT denoising trajectory where text conditioning shifts from global layout to fine-grained detail, then applies two steering mechanisms: Event-Partitioned Masking and Cross-Event Prompt Fusion. The authors also introduce Meve, a benchmark prompt suite for multi-event video generation, and report state-of-the-art results across 8 metrics with improved text alignment scaling with event count.
Researchers propose AGDO, a framework that replaces random masking in diffusion large language models (dLLMs) with attention-guided denoising order and token weighting during fine-tuning and reinforcement learning. The work is motivated by an empirical finding that tokens with stronger attention to unmasked context are more stable and critical for reasoning. Experiments on math and coding benchmarks show AGDO outperforms existing post-training methods for dLLMs, advancing the case for attention-aware training in parallel-decoding language models.
Researchers introduce AdaLook, an adaptive lookahead decoding framework for masked diffusion language models (DLMs) that dynamically determines rollout depth based on candidate-score variance rather than using fixed-depth lookahead. The method also enables branch expansion at informative intermediate states, avoiding unnecessary computation while improving exploration. Experiments across multiple benchmarks show AdaLook achieves a better accuracy-to-decoding-steps trade-off than existing one-step lookahead methods. The work addresses a known limitation of parallel text generation via DLMs, which are a non-autoregressive alternative to standard transformer decoding.
Researchers identify a small set of attention heads in vision-language model backbones, called 'gaze heads', whose attention patterns track the image region currently being described. Using comic strips as a controlled testbed, they show that intervening on the top-100 gaze heads (fewer than 9% of all heads) can steer the model to describe any chosen region at 83.1% accuracy, without retraining. The mechanism generalizes across model sizes from 2B to 32B parameters and to natural images (COCO), establishing a practical inference-time control lever for multimodal models via mechanistic analysis.
Researchers introduce SARDI, a training-free RAG framework for discrete diffusion language models that repurposes discarded low-confidence tokens during denoising as lookahead signals to guide retrieval before output is finalized. The method is retriever-agnostic and applicable to any reasoning-capable discrete diffusion LM. Evaluated across five multi-hop QA benchmarks, SARDI outperforms training-free diffusion and autoregressive retrieval baselines at up to 8x higher throughput.
ProtoAda is a new framework for Multimodal Continual Instruction Tuning (MCIT) that addresses a key failure mode in sparse Mixture-of-LoRA-Experts architectures: image-text similarity routing is format-blind and incorrectly merges tasks with similar semantics but different output structures (e.g., coordinate prediction vs. VQA). The method introduces format-aware task prototypes to guide both routing and adapter expansion, then consolidates compatible updates geometrically to reuse and refine existing parameters. Experiments across multiple benchmarks show improved performance, particularly on tasks whose answer formats are vulnerable to corruption by sequential fine-tuning.
This paper explores conditioning diffusion models on representations from pre-trained self-supervised models as an alternative to text prompts or semantic maps, which require large annotated datasets. The self-conditioning mechanism improves unconditional image generation quality and provides a controllable representation space. The authors identify directions of variation in this space and demonstrate smoothness and disentanglement properties, suggesting potential for fine-grained generative control without heavy annotation overhead.
Researchers at EPFL introduce Modus, a decoder-only model that treats all modalities symmetrically for any-to-any prediction without modality-specific heads, losses, or task pipelines. Unlike prior any-to-any models that use encoder-decoder or diffusion architectures trained from scratch, Modus leverages pre-trained decoder-only models as a prior. The model supports chained generation through intermediate modalities and cross-modal self-verification, achieving competitive performance with specialist and multitask baselines across benchmarks. All materials are open-sourced.