Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen† VILA Lab, MBZUAI † Corresponding author
Mohamed bin Zayed University of Artificial Intelligence
Real-time demo. Flash-dLLM (left) and Fast-dLLM (right) answer the same 8 GSM8K questions with LLaDA-1.5 on the same GPU, 4 questions per forward pass. Playback is fast-forwarded after Flash-dLLM finishes.
Real-time demo. Flash-dLLM (left) and LLaDA without a KV cache (right) answer the same 8 GSM8K questions on the same GPU, 4 questions per forward pass. Playback is fast-forwarded after Flash-dLLM finishes.

Diffusion large language models (dLLMs) generate text by iterative denoising, so they can decode many tokens at once. In practice their inference is still slow. Every denoising step revisits the whole sequence, which forces frequent KV-cache updates, and parallel decoding has to stay conservative to protect quality.

Flash-dLLM is a training-free framework that speeds up both parts together. Flash-Cache is an IO-aware KV cache built on a fused GPU kernel. Flash-Verify is a draft-and-verify decoder in which the dLLM checks its own drafts, with no auxiliary model. On LLaDA-1.5, Flash-dLLM reaches 210.6 tokens/s on GSM8K. It is 5.1× and 11.0× faster than Elastic-Cache on GSM8K and HumanEval, and up to 148× faster than decoding without a cache.

Two line plots against batch size from 1 to 32: throughput and peak GPU memory for LLaDA-1.5, Fast-dLLM, dLLM-Cache, Flash-dLLM and Llama3.
Figure 1. Throughput and peak GPU memory against batch size on GSM8K (generation length 512, LLaDA-1.5). Flash-dLLM scales to batch size 32 and needs less memory than the other dLLM methods. Fast-dLLM runs out of memory at batch size 24. Llama3-8B is an autoregressive reference.

The bottleneck: why a KV cache is not enough

KV caching and parallel decoding are usually studied separately. When the two are combined, three effects limit the speedup.

Three plots: latency of the conventional and the fused cache pipeline, the attention share of the most-attended tokens across layers, and average confidence against the number of early-decodable tokens.
Figure 2. The three observations behind Flash-dLLM. (a) The fused cache pipeline is 1.37× faster per layer. (b) A small set of decoded tokens receives most of the attention. (c) Higher confidence goes with more tokens that could be decoded early.

The Flash-dLLM approach

Flash-dLLM has two components. They share one fused kernel and one pre-allocated KV cache.

Overview diagram. Left: Flash-Cache, where tracked tokens and the masked window attend to the full KV cache. Right: Flash-Verify, where drafted and masked copies of the same positions are verified under a causal attention mask.
Figure 3. Overview of Flash-dLLM. Left, Flash-Cache: a fixed-size query of tracked tokens and the current masked window attends to the full KV cache. Right, Flash-Verify: low-confidence positions are duplicated as a draft view and a mask view, and a draft is accepted when both views agree.

Flash-Cache: IO-aware KV caching

The first pass processes the whole sequence and fills the cache. Every later step recomputes only a small query of fixed size and reuses the cache for everything else. Three pieces make this cheap.

Memory flow of Flash-Cache. Left: the fused kernel compared with the conventional KV-cache flow between GPU SRAM and HBM. Right: a block table that schedules attention for queries of different lengths.
Figure 4. IO-aware memory management in Flash-Cache. (a) The fused kernel keeps intermediate states in SRAM and writes keys and values directly to the cache. (b) A block table schedules attention for queries of different lengths.

Flash-Verify: the model verifies its own drafts

Flash-Verify recovers the correct tokens that a confidence threshold would throw away. The dLLM is both the drafter and the verifier.

  1. Draft. One forward pass predicts a token for every position in the masked window.
  2. Accept the confident tokens. Predictions with confidence of at least ε are decoded directly.
  3. Verify the rest in one extra pass. Each remaining position enters the query twice: once filled with its draft token (the draft view) and once as [MASK] (the mask view). A causal attention mask lets a mask-view position see the drafts before it in the order, but never its own draft.
  4. Accept until the first mismatch. A draft is accepted when the mask view predicts the same token with confidence of at least γ. As in speculative decoding, acceptance stops at the first rejection.

The verify pass reuses the fused kernel and the cache, and its cost grows with the masked window, not with the sequence. Flash-Verify decodes 5.7 tokens per step, against 2.8 for confidence-aware decoding and 1.0 for greedy decoding. A block of m accepted tokens deviates from the model's own sequential distribution by at most 1 − γm in total variation.

Accuracy, throughput and tokens per iteration for Flash-Verify and confidence-aware decoding at thresholds from 0.60 to 0.90.
Figure 5. Flash-Verify against confidence-aware decoding at different thresholds. Flash-Verify decodes more tokens per iteration and keeps a higher throughput at similar accuracy.

Results

All numbers are measured with LLaDA-1.5 on a single NVIDIA A100 80GB.

Four benchmarks at generation length 512

Flash-dLLM has the highest throughput on every benchmark. Its accuracy is higher than the uncached baseline on GSM8K, MATH and MBPP, and 0.61 points lower on HumanEval.

BenchmarkMetricLLaDA-1.5Fast-dLLMElastic-CacheFlash-dLLM
GSM8K (5-shot)TPS ↑2.6 (1.0×)36.8 (14.2×)41.7 (16.0×)210.6 (81.0×)
Score ↑81.3580.8282.7983.02
MATH (4-shot)TPS ↑5.0 (1.0×)44.4 (8.9×)41.4 (8.3×)210.1 (42.0×)
Score ↑35.6333.6835.8435.98
HumanEval (0-shot)TPS ↑3.2 (1.0×)15.4 (4.8×)16.8 (5.2×)185.6 (58.0×)
Score ↑40.8536.5937.8040.24
MBPP (3-shot)TPS ↑1.0 (1.0×)17.8 (17.8×)32.8 (32.8×)148.2 (148.2×)
Score ↑38.2036.2039.0039.00

TPS is throughput in tokens per second, with the speedup over LLaDA-1.5 decoded greedily without a cache. Score is accuracy or pass@1 in percent. Flash-dLLM is Flash-Cache with Flash-Verify.

Comparison with more acceleration methods

Method (GSM8K, length 512)TPS ↑Score ↑
dKV-Cache14.9 (5.7×)81.50
FlashDLM15.7 (6.0×)79.91
dLLM-Cache16.8 (6.5×)80.97
Fast-dLLM36.8 (14.2×)80.82
Dyna-dLLM38.4 (14.8×)79.32
Elastic-Cache41.7 (16.0×)82.79
FreeDave42.8 (16.5×)80.97
Flash-Cache (confidence-aware)149.4 (57.5×)82.87
Flash-Cache + Flash-Verify210.6 (81.0×)83.02

Scaling with batch size

Throughput keeps growing up to batch size 32. At batch size 16, Flash-dLLM uses about 26 GB of GPU memory, against about 50 GB for Fast-dLLM (Figure 1).

Flash-Cache withTokens / stepTPS at batch size
1481632
Greedy decoding1.018.537.746.351.355.0
Confidence-aware decoding2.851.7102.4120.5131.8139.5
Flash-Verify5.756.0131.3164.8186.2199.8

GSM8K, 5-shot, generation length 512.

Accuracy and speed can be traded

The verification threshold γ and the tracking budget set the operating point. Across the tested settings, Flash-dLLM moves from 277.8 tokens/s at 80.23% accuracy to 131.6 tokens/s at 83.41%.

Two plots of average throughput against average accuracy on GSM8K: the Pareto frontier of Flash-Verify, and the frontiers of Flash-Verify and confidence-aware decoding.
Figure 6. Accuracy against throughput on GSM8K for different tracking budgets and thresholds, averaged over five seeds. Flash-Verify is faster than confidence-aware decoding over most of the accuracy range they share.

The same answer in fewer steps

One GSM8K prompt, decoded by LLaDA-1.5 without a cache. All three decoders return the correct answer.

DecodingSteps ↓TokensTime ↓
Greedy28328364.7 s
Confidence-aware6627717.2 s
Flash-Verify3027613.1 s

Conclusion

Flash-dLLM treats GPU memory traffic and token verification as one problem. A fused kernel makes the KV cache cheap to keep fresh, and the same cache lets the model verify its own drafts. Decoding gets faster and uses less memory, with no retraining and no second model.

Citation

@misc{nguyentri2026flashdllm,
  title         = {{Flash-dLLM}: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs},
  author        = {Nguyen-Tri, Quan and Ranjan, Mukul and Shen, Zhiqiang},
  year          = {2026},
  eprint        = {2609.26796},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}