MATH · IN · MODELS

A real trained graph transformer's attention-weighted effective graph undergoes curvature collapse, concentrating on and worsening bottleneck edges rather than smoothing them

measured in 1 paper

Balanced Forman Curvature is computed on both the raw input graph and an attention-re-weighted 'effective graph' formed from real trained Graph Transformer (GT, GraphiT, SAN) attention weights, across ZINC, Tox21, and Long-Range Graph Benchmark peptide datasets [tori-etal-2026-probing-gnn-activation-patterns-graph-topology] On LRGB peptide graphs, the attention-weighted effective graph undergoes a Curvature Collapse: the fraction of negatively-curved edges jumps from 57% to 84% on peptides-func and from 57% to 82% on peptides-struct, with the weighted Balanced Forman Curvature shifting from -0.68 to -0.70 and the spectral gap dropping [tori-etal-2026-probing-gnn-activation-patterns-graph-topology] Attention concentrates on and worsens existing topological bottlenecks rather than smoothing them, contrary to the theoretical prediction that massive attention weights would preferentially target curvature-extreme edges to counteract them; causal pruning of massive-activation edges within bottlenecks spikes loss by 22-27%, confirming the effect is functionally load-bearing rather than an artifact [tori-etal-2026-probing-gnn-activation-patterns-graph-topology]

Context

Balanced Forman Curvature, massive activations, graph bottlenecks and over-squashing, causal edge pruning

Papers

Probing Graph Neural Network Activation Patterns Through Graph Topology — Tori, Floriano, Bini, Lorenzo, Sorbi, Marco, Marchand-Maillet, Stephane, Ginis, Vincent2026 · arXiv:2602.21092