ADP 前沿学习

← 板块一 · 研究前沿

Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

Kim · cs.LG,cond-mat.stat-mech,stat.ML · 2026-09-03 · 原文 · 证据态 未学

We derive two attention operators from generalized statistical entropies. Kaniadakis entropy yields an exact full-support normalization whose weights and low-score sensitivities decay algebraically, rather than exponentially as in Softmax or by exact truncation as in entmax. Classical Abe entropy yields an implicit reciprocal-symmetric operator. With q=e^ε, the involution q\leftrightarrow q^{-1} removes every odd correction about Softmax; we obtain the normalized second- and fourth-order terms, including the deformation of the normalization multiplier. These stationary laws follow from a Fisher-metric Lagrangian on the probability simplex, whose Shannon sector recovers scaled dot-product Softmax. We also give a tangent-gradient test for deciding whether changing the entropy changes the attention profile or only its scale. Rényi and two-parameter Sharma--Mittal entropies retain the Tsallis--entmax inverse-gradient shape, but their global moments make the effective temperature input dependent when the external temperature is fixed. Distinguishing profile-shape equivalence from fixed-parameter operator equivalence separates new normalization shapes from adaptive rescalings and org

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

We derive two attention operators from generalized statistical entropies.

Kaniadakis entropy yields an exact full-support normalization whose weights and low-score sensitivities decay algebraically, rather than exponentially as in Softmax or by exact truncation as in entmax.

2. 领域脉络

本文类目:cs.LG、cond-mat.stat-mech、stat.ML,属于其所在研究脉络的最新进展。

3. 机制拆解

Classical Abe entropy yields an implicit reciprocal-symmetric operator.

These stationary laws follow from a Fisher-metric Lagrangian on the probability simplex, whose Shannon sector recovers scaled dot-product Softmax.

4. 证据与数字

With q=e^ε, the involution q\leftrightarrow q^{-1} removes every odd correction about Softmax; we obtain the normalized second- and fourth-order terms, including the deformation of the normalization multiplier.

5. 反例与边界

We also give a tangent-gradient test for deciding whether changing the entropy changes the attention profile or only its scale.

Rényi and two-parameter Sharma--Mittal entropies retain the Tsallis--entmax inverse-gradient shape, but their global moments make the effective temperature input dependent when the external temperature is fixed.

6. 跨领域连接与意外收获

横跨 3 个类目(cs.LG、cond-mat.stat-mech、stat.ML),关注其在你兴趣板块间的迁移面。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。

主动回忆

先合上内容自己复述,再点「显示」核对,然后如实自评。评分即时进 FSRS 排程。