Remark. What the attention-twin experiment observes [ftip-00BY]
Remark. What the attention-twin experiment observes [ftip-00BY]
Section 6 and Table 3 of The loss does not see the basis, but Adam does[singh2026lossbasis] compare two function-equivalent initializations of a two-layer, four-head transformer on modular addition. The reported relative logit distance after one Adam step is \(3.6\times 10^{-3}\), versus \(2.9\times 10^{-7}\) for a same-basis noise twin; the final per-head invariant distance is reported as 56 percent.
The exact first-step defect has an algebraic explanation, but later drift, scale robustness, and downstream behavior are empirical. The tested small transformers, precision, optimizer allocation, seeds, and task do not establish a frontier-wide law.