Attention-graph Laplacian spectra separate valid from invalid proofs
measured in 1 paper- Treating each layer's attention matrix as a weighted graph, four training-free Laplacian spectral diagnostics (Fiedler value lambda2, high-frequency energy ratio HFER, spectral entropy and smoothness) separate valid from invalid mathematical proofs. [noel-2026-geometry-of-reason] - The separation is large and highly significant across all seven models (Cohen's d 2.09-3.30; per-model Mann-Whitney p from 1.16e-48 to 4.51e-66; peak d=3.30 for Phi-3.5-mini, with a pooled significance p<1e-116). [noel-2026-geometry-of-reason] - A single calibrated threshold on one spectral metric classifies proofs at 85.0-95.6% accuracy (93-95% on the full dataset), dropping to 82.8-85.9% under nested cross-validation. [noel-2026-geometry-of-reason] - Global-attention models (Llama, Qwen, Phi) carry the signal in HFER, whereas Mistral-7B's Sliding-Window Attention shifts it to late-layer smoothness (d=2.09, p=1.16e-48). [noel-2026-geometry-of-reason] - Some proofs the spectral method flags as valid are rejected by Lean/Isabelle only for technical reasons (timeouts, missing imports), which the paper terms "Platonic validity". [noel-2026-geometry-of-reason] - Tested on MiniF2F proofs across seven models / four families (Llama-3.2-1B/3B, Llama-3.1-8B, Qwen2.5-0.5B/7B, Phi-3.5-mini, Mistral-7B-v0.1 — named as base checkpoints though the text calls them instruction-tuned); observational. [noel-2026-geometry-of-reason]