Transfer Learning in High-Dimensional Clustering: Minimax Thresholds and Applications in Single-Cell Data

Abhinav Chakraborty, Sagnik Nandy

Abstract

Clustering is a fundamental problem in statistics, with applications across many scientific disciplines. In many modern applications involving clustering, the primary dataset (the target data) is accompanied by related datasets (the source data). Transferring information from such sources may improve clustering accuracy in the target, making transfer learning for clustering practically important. Despite recent progress, the conditions under which source data improve target clustering remain unclear in high-dimensional settings, even for the canonical Gaussian mixture model. In this paper, we study the clustering problem in a two-community Gaussian mixture model where relatedness is captured by the geometric alignment of the target and source cluster means. We develop a minimax-optimal transfer-assisted clustering procedure and characterize, up to logarithmic factors, the phase transition for consistent target clustering in terms of the signal-to-noise ratios, sample sizes, ambient dimension, and degree of alignment between the datasets. The technique is also extended to adaptively choose between the target-only or the source assisted clustering depending on the target signal strength. Furthermore, we also extend our techniques to accommodate multiple communities and and multiple source datasets. Extensive simulations and an analysis of a human lung single-cell RNA-sequencing atlas demonstrate the practical effectiveness of our methods.

Disclosure

“plore such areas in future research. Use of generative AI. The authors have used ChatGPT 5.5 and 5.6 for proof checking, copyediting text and brainstorming ideas. The major ideas underlying some proofs also developed from discussion with ChatGPT. The authors however verified correctness of all results, re-wrote the arguments and own complete responsibility. The authors have also used Claude code for simulations and data analysis. Once again, all codes were carefully checked by the”

PDF page 23
Classification
Proof ideas or individual proof-step assistance
Multiplier
8
Verified

Structural counts

Pages 90 pdf
Theorems 10 source
Lemmas 30 source
Propositions 0 source
Corollaries 0 source
Definitions 0 source
Displayed equations 776 source
Bibliography entries 123 source
Appendix pages 0 estimated

Count notes

  • Source counts use the expanded primary TeX file arxiv_version.tex.
  • Appendix pages include the first PDF page with an explicit Appendix heading through the final page.