Soft-masking accelerates convergence in Masked Diffusion Language Models; the paper argues that linear interpolation in the embedding space is flawed due to near-constant token/mask embedding angles during training.
Read the original at arxiv.org→arXiv:2608.06529v1 Announce Type: new Abstract: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (LERP) in the...
Original headline: "Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.CL lead source Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models
- Aug 10, 04:00 UTC arXiv cs.LG Retrofitting Linear Attention into Diffusion Language Models