MATH · IN · MODELS

A 141-hypothesis automated topology/geometry audit of scGPT and Geneformer gene-embedding representations finds statistically significant persistent homology in nearly every transformer layer and a canonical-correlation cross-model alignment of 0.80 between the two independently-trained models

measured in 1 paper

Kendiukhov (2026) runs an autonomous executor-brainstormer loop that proposed, tested, and refined 141 geometric/topological hypotheses across 52 iterations about gene-embedding representations in scGPT and Geneformer, with explicit null controls and disjoint gene-pool splits. Persistent homology is statistically significant (p<0.05) in 11/12 transformer layers in the weakest domain and 12/12 in the other two; manifold-aware distance metrics outperform Euclidean distance for identifying regulatory gene pairs; graph community partitions track known transcription-factor-target relationships; canonical correlation analysis between scGPT and Geneformer's independently- trained gene-embedding spaces yields a canonical correlation of 0.80 and 72% gene-retrieval accuracy (though none of 19 tested methods reliably recover exact gene-level correspondence); robust signal concentrates in immune tissue under stringent nulls.

Context

genomics, persistent-homology, cross-model-alignment

Papers

What Topological and Geometric Structure Do Biological Foundation Models Learn? Evidence from 141 Hypotheses — Kendiukhov, Ihor2026 · arXiv:2602.22289