Theorem. DGG finite-batch head-gradient inequality [ftip-0081]

For \(T\geq 1\) active, unclipped tokens with fixed sampled histories and the untied output head of Theorem [ftip-007Z], let \[ G^{\rm lm}=\frac 1T\sum _{i=1}^T G_i^{\rm lm}, \qquad \overline {r^2}=\frac 1T\sum _{i=1}^T r_i^2, \] and set \[ c_{\max }= \max _{1\leq i\leq T} \widehat A_i^2 \|e_{a_i}-\pi _\theta (\mathord \cdot \mid h_{L,i})\|_2^2 \|h_{L,i}\|_2^2. \] Then \[ \|G^{\rm lm}\|_F^2\leq c_{\max }\overline {r^2} =c_{\max }(1+\widehat \chi ^2), \] where \(\widehat \chi ^2\) is the statistic of Definition [ftip-0066]. If \(c_{\max }>0\), rearrangement gives the source's equivalent direction \[ \widehat \chi ^2\geq \frac {\|G^{\rm lm}\|_F^2}{c_{\max }}-1. \]

The inequality points from observed head-gradient energy to a lower bound on this finite-batch squared-ratio statistic. It is not a converse: cancellation can hide nonunit ratios. The active cancellation example in Example [ftip-0084] states its clipping interval explicitly. The proof uses no intermediate-weight bound.

The finite-batch inequality follows the calculation in [miao2026when, Section 4.2.2, Lemma 2, Theorem 2, and Appendix B], under the untied-head and fixed-history assumptions stated above.