Back to feed
arXiv cs.CL
arXiv cs.CL
8/5/2026
Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

Short summary

This paper introduces speculative correction for diffusion language models (DLMs): a small model drafts a full response, then a larger model refines it via bidirectional diffusion. Using LLaDA2.1-Flash and Mini, the approach improves GSM8K accuracy from 0.848 to 0.899 while running 1.2x faster, and MBPP from 0.545 to 0.693. The Mini-Flash cascade achieves near-Flash quality on MATH while running 2.17x faster, offering a training-free route to faster DLM generation.

  • Speculative correction: small DLM drafts, larger DLM refines via bidirectional diffusion
  • Flash-Flash self-refinement boosts GSM8K 0.848→0.899 at 1.2x speed; MBPP 0.545→0.693
  • Mini-Flash cascade matches Flash MATH quality at 2.17x speedup, no additional training needed

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more