Shiva-DiT: Residual-Based Differentiable Top-k Selection for Efficient Diffusion Transformers
Zhang, Zhao, Wu, Sun, Zhao, Deng · cs.LG,cs.AI,cs.CV · 2026-09-02 · 原文 · 证据态 未学
Diffusion Transformers (DiTs) are costly at high resolution because self-attention scales quadratically with token sequence length. Existing pruning methods do not jointly provide end-to-end learnability, low training overhead, and deterministic token counts for predictable token-dependent computation. We propose Shiva-DiT, based on Residual-Based Differentiable Top-k Selection. Its forward pass executes hard top-k selection, while a residual-aware straight-through estimator propagates gradients to both token scores and the budget k without evaluating a second backbone path. A Context-Aware Router and Adaptive Ratio Policy learn layer- and timestep-dependent retention schedules under a target average budget. Experiments on SD3-Medium, Flux.1-dev, and PixArt-Σ show consistent reductions in FLOPs and measured latency. On SD3-Medium, Shiva-DiT provides four fidelity-latency operating points and reaches a 1.54x wall-clock speedup with competitive fidelity.
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
Diffusion Transformers (DiTs) are costly at high resolution because self-attention scales quadratically with token sequence length.
Existing pruning methods do not jointly provide end-to-end learnability, low training overhead, and deterministic token counts for predictable token-dependent computation.
2. 领域脉络
本文类目:cs.LG、cs.AI、cs.CV,属于其所在研究脉络的最新进展。
3. 机制拆解
We propose Shiva-DiT, based on Residual-Based Differentiable Top-k Selection.
Its forward pass executes hard top-k selection, while a residual-aware straight-through estimator propagates gradients to both token scores and the budget k without evaluating a second backbone path.
A Context-Aware Router and Adaptive Ratio Policy learn layer- and timestep-dependent retention schedules under a target average budget.
4. 证据与数字
Experiments on SD3-Medium, Flux.1-dev, and PixArt-Σ show consistent reductions in FLOPs and measured latency.
On SD3-Medium, Shiva-DiT provides four fidelity-latency operating points and reaches a 1.54x wall-clock speedup with competitive fidelity.
5. 反例与边界
摘要未声明局限与反例——这是需要警惕的信号,精读时先问边界。
6. 跨领域连接与意外收获
横跨 3 个类目(cs.LG、cs.AI、cs.CV),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。
主动回忆
先合上内容自己复述,再点「显示」核对,然后如实自评。评分即时进 FSRS 排程。
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
Diffusion Transformers (DiTs) are costly at high resolution because self-attention scales quadratically with token sequence length.
Existing pruning methods do not jointly provide end-to-end learnability, low training overhead, and deterministic token counts for predictable token-dependent computation.
2. 领域脉络
本文类目:cs.LG、cs.AI、cs.CV,属于其所在研究脉络的最新进展。
3. 机制拆解
We propose Shiva-DiT, based on Residual-Based Differentiable Top-k Selection.
Its forward pass executes hard top-k selection, while a residual-aware straight-through estimator propagates gradients to both token scores and the budget k without evaluating a second backbone path.
A Context-Aware Router and Adaptive Ratio Policy learn layer- and timestep-dependent retention schedules under a target average budget.
4. 证据与数字
Experiments on SD3-Medium, Flux.1-dev, and PixArt-Σ show consistent reductions in FLOPs and measured latency.
On SD3-Medium, Shiva-DiT provides four fidelity-latency operating points and reaches a 1.54x wall-clock speedup with competitive fidelity.
5. 反例与边界
摘要未声明局限与反例——这是需要警惕的信号,精读时先问边界。
6. 跨领域连接与意外收获
横跨 3 个类目(cs.LG、cs.AI、cs.CV),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。