U-centering as subset ANOVA: edge regression and higher-order theory

Xianyang Zhang

Abstract

The unbiased sample versions of squared distance covariance and the Hilbert-Schmidt independence criterion (HSIC) are fourth-order U-statistics, yet U-centering evaluates them from pairwise arrays in $O(n^2)$ operations. We show that U-centering is exactly the least-squares residual obtained after fitting additive endpoint effects to a symmetric hollow array. This interpretation explains the zero row sums and the denominator $n(n-3)$ through the residual degrees of freedom. The same pairwise residualization also gives useful regression identities. After endpoint effects are removed from both arrays, the U-centered dependence $t$-statistic is the ordinary slope $t$-statistic obtained by regressing one adjusted array on the other. In the two-sample problem, pooling the observations and using the between-group pair indicator as the predictor shows that the generalized-energy statistic is twice the fitted slope. The common-endpoint and fully interacted regressions give the same slope but use different residual standard errors. For $n\ge2r$, we extend the construction to arrays indexed by $r$-subsets. Higher-order U-centering removes all effects involving fewer than $r$ sample labels, leaves zero $(r-1)$-way margins, and projects onto a residual space of dimension $\binom nr-\binom n{r-1}$. For two symmetric kernels with $r$ arguments, the normalized inner product of the centered arrays is unbiased for the cross-moment of their $r$th Hoeffding components. A direct estimator can involve products spanning as many as $2r$ observations, but subset-margin inversion or higher-order U-centering evaluates the same quantity in $O(n^r)$ operations for fixed $r$. When both arrays are formed from the same kernel and sample, this becomes a nonnegative unbiased estimator of the variance of the highest-order Hoeffding component.

Disclosure

“AI-use statement During the preparation of this manuscript, the author used GPT-5.6 Sol for language editing and exploratory assistance with proof development. The author independently checked every definition, statement, and proof and takes full responsibility for the content of the manuscript. References [1] Blozneli”

PDF page 35
Classification
Proof ideas or individual proof-step assistance
Multiplier
8
Verified

Structural counts

Pages 36 pdf
Theorems 5 source
Lemmas 1 source
Propositions 6 source
Corollaries 3 source
Definitions 2 source
Displayed equations 225 source
Bibliography entries 24 source
Appendix pages 17 estimated

Count notes

  • Source counts use the expanded primary TeX file main.tex.
  • Appendix pages include the first PDF page with an explicit Appendix heading through the final page.