Theorem. DGG occurrence-to-head gradient-energy bound [ftip-0080]

Use the per-token quantities and untied head of Theorem [ftip-007Z], and fix one occurrence \(o\in \mathcal O_i\). Assume \(d_{\rm model}\geq 1\) and positive constants \(\alpha _{\min }\), \(\beta _{\max }\), and \(C\) satisfy \[ \|h_{L,i}\|_2^2\geq \alpha _{\min }d_{\rm model}, \qquad \|x_o\|_2^2\leq \beta _{\max }d_{\rm model}, \] and the two logit-sensitivity bounds \[ \mathbb E_{a\sim p_i} \|(J_{io})_{a,:}\|_2^2\leq C, \qquad \|(J_{io})_{a_i,:}\|_2^2\leq C. \] If \(\widehat A_i\neq 0\) and \(p_i(a_i)<1\), then \[ \frac {\|H_{io}\|_F^2}{\|G_i^{\rm lm}\|_F^2} \leq \frac {\mathcal C_{\rm occ}} {(1-p_i(a_i))^2}, \qquad \mathcal C_{\rm occ} =\frac {4\beta _{\max }C}{\alpha _{\min }}. \]

This bounds one occurrence contribution, not the total derivative of a shared intermediate weight. For the latter, the sum in Theorem [ftip-007Z] gives only \[ \|G_i^{\rm int}\|_F \leq \sum _{o\in \mathcal O_i} \|J_{io}^{\mathsf T}E_i\|_2\|x_o\|_2. \] If the same displayed hypotheses hold for every occurrence, the triangle inequality yields the ratio bound \(|\mathcal O_i|^2\mathcal C_{\rm occ}/(1-p_i(a_i))^2\) for \(\|G_i^{\rm int}\|_F^2/\|G_i^{\rm lm}\|_F^2\). Controlling only the application at position \(i\) does not establish this all-occurrence hypothesis or the source's claimed bound for the total gradient. The contributing positions, reuse pattern, and Jacobians depend on the architecture. Neither bound establishes a small gradient or an evaluation-score comparison.

The activation and row-energy hypotheses and the local inequality follow the calculation in [miao2026when, Section 4.2.1, Lemma 1, Assumption 1, Theorem 1, and Appendix A.3], with the differentiated object restricted to an occurrence contribution. They do not prove the source's total shared-weight claim under its local hypotheses.