MATH · IN · MODELS
structures / Manifolds / Intrinsic-dimension profile across depth

Intrinsic-dimension profile across depth

PROPERTYmeasurementfunctionalintermediatehow it's classified →

The dimensionality of the data manifold a network's own representations lie on is not fixed — it expands sharply in early layers, then contracts to a low-dimensional plateau or minimum, and the layer at that minimum tends to carry the most abstract, task-relevant semantic content.

Replicationcomputed from the corpus — never hand-assigned
38 papers · no shared authors11 architecture classes · across papers8 domains · across papers42 model families · across papers
Filled = two or more values reported by papers that share no author — replication. Outlined = two or more values, but all from a single study — breadth, not replication. Grey = a single value. Derived from paper authorship and each model's architecture class, domain and family; it updates itself when a paper is added.

Definition

Rather than describing one manifold’s fixed shape, this concept describes how the manifold’s own intrinsic dimension (ID) — the minimum number of coordinates needed to describe the data locally without loss, measured here via the TwoNN nearest-neighbor estimator — changes as a function of network depth. Across the self-supervised transformers studied so far, the same qualitative trajectory recurs: an early expansion phase (ID rises sharply, the neighbor structure churns rapidly layer-to-layer), a compression phase (ID falls and stabilizes at a much lower plateau or local minimum, with high layer-to-layer neighbor stability), and — in some models — a final re-expansion/decoding phase as the representation is unpacked back toward the reconstruction target.

Relative to a single fixed manifold

Every other node under Manifolds (circle, torus, sphere, paraboloid, …) describes one manifold’s shape at a fixed site. This concept instead tracks a sequence of manifolds, one per layer, and treats their changing dimensionality itself as the object of interest — an encoder/decoder-style trajectory through representation space, rather than a static geometric claim about any one layer.

Key evidence

Valeriani, Doimo, Cuturello, Laio, Ansuini & Cazzaniga (2023) measure ID (Intrinsic dimension estimation (TwoNN)) and layer-to-layer neighbor stability (Neighborhood overlap) across ESM-2 protein language models (35M/650M/3B) and iGPT image transformers (S/M/L, 76M/455M/1.4B), finding the same expansion-then-compression shape in every model, with the ID’s local minimum in each case coinciding with the layer where neighborhood overlap with ground-truth semantic labels (remote protein homology; ImageNet class) peaks — i.e. the geometrically most-compressed layer is also the most semantically abstract one, in every model tested. A brief preliminary appendix extends this to Llama-2-70B on sentiment classification (SST): the ID profile is more complex (three peaks, two local minima) than in either protein or image models, but the qualitative link holds — overlap with the sentiment label partition is highest at the first local ID minimum. See Relation frame (ordered multi-token tuple geometry) and Separability/alignment decomposition for other depth-wise geometric decompositions in this corpus, on different structural objects (relation tuples; separability/alignment) than raw manifold dimensionality.

Extension to generative-latent and molecular-embedding manifolds

Choi, Hwang, Cho & Kang (2023) extend this concept from a per-layer scalar profile to a spatially-varying field over a single real trained GAN’s (StyleGAN/StyleGAN2) latent manifold: local intrinsic dimension is estimated at each latent point via a Riemannian-manifold treatment of the generator, and correlates strongly with supervised disentanglement scores despite requiring no attribute labels. See choi-etal-2023-local-intrinsic-dimension-of-a-real-stylegan2-latent-manifold-varies-across-latent-space-and-correlates-with-disentanglement. El-Samman, Husain, Huynh, De Castro, Morton & De Baerdemacker (2024) measure a related global-effective-dimension finding on a real trained SchNet-family GNN’s molecular embeddings (QM9): the fully-trained 128-parameter embedding space reduces to roughly 5 effective parameters via dimension reduction, alongside a linear-separability finding on the same embeddings (see Linear Separability). See el-samman-etal-2024-a-real-schnet-family-gnns-128-parameter-qm9-molecular-embedding-space-reduces-to-around-5-effective-parameters-with-sharp-linear-boundaries-separating-chemical-moieties.

Extension to brain alignment, and a causal fine-tuning intervention

Cheng, Vaidya & Antonello (2026) estimate per-layer intrinsic dimension (via GRIDE) in real OPT (125M-13B), Pythia (160M-6.9B), WavLM (base-plus, large) and Whisper (large) activations, finding layerwise ID correlates strongly with how well that layer’s activations predict real human brain responses (fMRI rho=0.76; ECoG rho=0.43), with the ID peak layer and the best-brain-predicting layer usually 0-1 layers apart. Causal validation: fine-tuning the best-performing WavLM layer to better predict real fMRI responses (“brain-tuning”) causally raises both encoding performance and that layer’s own measured intrinsic dimension, with a random-Fourier-features negative control showing raised ID alone does not suffice to cause the effect — a genuine causal manipulation of the ID profile itself, distinct from the purely observational depth-wise measurements elsewhere in this node. See layerwise-intrinsic-dimension-of-real-opt-pythia-wavlm-and-whisper-activations-peaks-near-the-layer-that-best-predicts-real-fmri-ecog-brain-responses-and-brain-tuning-a-layer-causally-raises-both-its-id-and-its-brain-alignment.

Extension to a non-language, non-vision domain: single-cell transcriptomics

Kendiukhov (2026) extends the “compression coincides with peak semantic content” pattern to a genuinely new domain and a new spectral measure: across scGPT’s 12 transformer layers, the SVD effective rank of the gene-embedding matrix collapses monotonically 23.6 to 1.6 (Spearman rho=-1.000), with the top singular vector’s variance fraction rising from 53.7% to 93.4% and TwoNN intrinsic dimensionality falling 32.6 to 18.1. Unlike the ESM-2/iGPT case (a local ID minimum partway through the network, followed by re-expansion), this is a monotonic collapse to the final layer — but the same qualitative claim holds: the most spectrally-compressed layer’s low-rank subspace is the one that best correlates with independent ground truth (STRING protein-protein- interaction confidence, TRRUST transcription-factor-target relations, cell-type identity), with a feature-shuffle control confirming the collapse is not an artifact. See scgpt-gene-embedding-effective-rank-collapses-14-fold-across-layers-and-the-final-compressed-layer-encodes-ppi-tf-target-and-cell-type-ground-truth.

Extension to whole-sequence shape space, rather than pointwise activations

Beshkov & Malthe-Sørenssen (2026) apply a categorically different notion of dimensionality to the same expansion-then-compression question: rather than TwoNN on pointwise activations, they treat each protein’s whole per-residue PLM representation as a curve, map it to its square-root- velocity (SRV) shape via SRV shape-space Fréchet radius and tangent-PCA effective dimension, and measure the tangent-PCA effective dimension of the resulting shape space across layers of ESM2 (35M/150M/650M/3B) and Ankh. The same qualitative expansion-then-compression trajectory recurs — larger models expand the shape-space dimension more sharply in early layers before contracting — but at a far lower absolute dimension than PCA run directly on the same models’ flattened pointwise activations, indicating that while individual residue activations occupy a high-dimensional ambient space, the ways in which different protein shapes (whole trajectories) differ from each other are describable by only a handful of directions. A companion graph-filtration analysis on the same models finds three-dimensional protein structure is most faithfully encoded at very short (2-residue) and moderate (8-residue) context lengths, and closer to but before the final layer — independent evidence that the most structurally-faithful layer is not the last one. See beshkov-malthe-sorenssen-2026-plm-shape-spaces-undergo-the-same-expansion-then-compression-effective-dimension-trajectory-as-pointwise-activations-but-at-a-far-lower-absolute-dimension.

A conflicting depth trend, not yet reconciled

Cai, Huang, Bian & Church (2021) estimate a related but distinct local-dimension statistic — Local Intrinsic Dimension (LID), via a KK-nearest-neighbor expansion-model estimator rather than TwoNN — across BERT, DistilBERT, GPT, GPT-2, and ELMo, and find LID increases nearly linearly with depth in every model, the opposite trajectory shape from the expansion-then-compression pattern documented above. Both findings stand as verified in this map; the discrepancy may reflect a genuine difference between language models and the protein/image models studied here, a genuine difference between the LID and TwoNN estimators, or both — flagged as an open question rather than resolved. See manifold-dimension-lower-than-ambient.

A theoretical impossibility casts doubt on the expansion-then-compression shape itself

Schulte & Rügamer (2026) prove that because standard network layers (linear/conv, ReLU, softmax, pooling, residual connections, BatchNorm/RMSNorm, and their compositions including self-attention) are Lipschitz maps, both pointwise and Hausdorff intrinsic dimension can only stay the same or decrease from one layer to the next — they can never rise. Yet reproducing the standard TwoNN/MLE/Gride pipeline on ResNet-34 (ImageNet) and on Llama-3.1-8B, Mistral-7B-v0.3, and Pythia-6.9B (WikiText prompts, last-hidden-state per layer) recovers the familiar hump-shaped early-expansion-then-compression pattern documented elsewhere in this node — a pattern their own theorem says is mathematically impossible for the true ID to exhibit. Rather than proposing a replacement estimator, they investigate what these estimators are actually tracking instead (nearest-neighbor distances, ambient dimension, cosine similarity, representation norm, entropy), concluding that the widely-cited “abstraction emerges as a mid-layer dimensionality peak” narrative built on TwoNN-style estimators is very likely an artifact of estimator bias rather than a real geometric signal. This is a direct methodological caution for every expansion-then-compression finding in this node that relies on TwoNN or closely related nearest-neighbor-ratio estimators (though not for findings, such as MST-based or spectral-entropy/effective-rank measures, that use different machinery). See lipschitz-layers-cannot-increase-true-intrinsic-dimension-so-the-standard-twonn-hump-shaped-id-profile-is-likely-an-estimator-artifact.

Relative to tangent-aligned-anisotropy

Tangent-Aligned Anisotropy Hypothesis proposes a mechanism for why representations end up occupying a low-dimensional subspace in the first place — frequency-concentrated sampling and self-reinforcing tangent-aligned gradients — rather than measuring the resulting dimensionality profile directly the way this concept does. The two are complementary: this concept characterizes what the dimensionality trajectory looks like across depth; that hypothesis proposes why training would produce low-dimensional, anisotropic representations at all.

Relative to sufficiency-staging

Attention–MLP Sufficiency Staging Hypothesis proposes a sharper, mechanistically-located version of the same kind of compression trajectory: rather than a generic expansion-then-plateau across many layers, it claims a two-stage split tied to one specific sublayer boundary (attention output vs. post-MLP), where what’s being compressed is a known, analytically-derivable sufficient statistic rather than “intrinsic dimension” in the abstract. This concept describes the general shape of the trajectory; that hypothesis describes one specific, more tightly-scoped mechanism producing a compression step within it.

Token-level local ID as a per-token, rather than per-layer, statistic

Lee, Weber, Viegas & Wattenberg (2025) apply a k-NN-neighborhood, PCA-based local intrinsic-dimension estimator (components needed for 95% variance) to individual tokens’ embeddings — not layer-by-layer activations of a shared input — across GPT2, Llama3, Gemma2, GPT-NeoX-20B and OLMo-7B, finding that low-ID tokens (GPT2-medium range 508-635, versus a random-Gaussian baseline of ~605-611) form semantically coherent clusters while high-ID tokens do not. This is a within-layer, across-vocabulary application of the same local-dimension-estimation toolkit used elsewhere in this map for across-depth profiles, showing the estimator also picks out meaningful structure when applied to a fixed layer’s full token population rather than a fixed input’s trajectory across layers. See lee-etal-2025-cross-model-embedding-orientation-similarity-drops-across-families-local-id-clusters-tokens-linear-map-transfers-steering-vectors.

Per-concept local ID as a complexity proxy, rather than a per-token or per-layer statistic

Skierś, Trzciński & Deja (2026, ELROND) apply local intrinsic-dimension estimation to a third axis beyond depth or token identity: individual concepts. Backpropagating the differences between stochastic image realizations of the same text prompt in real SDXL, then decomposing the resulting gradient directions via PCA or a sparse autoencoder, they estimate the LID of each decomposed concept-direction’s own manifold within the text embedding and find general concepts (e.g. “Dog”) consistently have higher LID than specific hyponyms (e.g. “Poodle”), validated against WordNet. Causally, injecting these directions into a distilled, mode-collapsed student model (SDXL-DMD) restores output diversity toward the teacher’s distribution (measured via FID), tying the per-concept dimensionality measurement to a genuine generative-diversity effect. See gradient-derived-concept-directions-in-real-sdxl-and-sdxl-dmd-text-embeddings-have-local-intrinsic-dimension-tracking-concept-generality-and-injecting-them-causally-restores-fid-lost-to-distillation.

A training-dynamics phase transition dissociating linear from nonlinear dimension

Lee, Jiralerspong, Yu, Bengio & Cheng (2024) track nonlinear intrinsic dimension (TwoNN) and linear PCA effective dimension across pretraining on Pythia-410M/1.4B/6.9B, finding nonlinear intrinsic dimension undergoes a sharp phase transition at training step t~10^3 that coincides with the onset of zero-shot task competence, consistent across model sizes and generalizing to fully-trained Llama-3-8B and Mistral-7B. Linear PCA effective dimension, by contrast, correlates with superficial/Kolmogorov data complexity (gzip compressibility) rather than with this emergence event — a dissociation between the linear and nonlinear dimensionality measures themselves, orthogonal to (and a useful complement to) this node’s own expansion-then-compression depth trajectory, since it concerns training dynamics at fixed depth rather than a depth trajectory at fixed training time. See lee-etal-2024-nonlinear-intrinsic-dimension-shows-a-sharp-phase-transition-coinciding-with-zero-shot-task-competence-while-linear-pca-dimension-tracks-only-superficial-complexity.

Local intrinsic dimension predicts LLM truthfulness with a hunchback layer profile

Yin, Srinivasa & Chang (2024) compute Local Intrinsic Dimension (LID, via GeoMLE) on the per-layer activations of real Llama-2-7B and Llama-2-13B generating answers on four real QA datasets (TriviaQA, HotpotQA, TydiQA-GP, CoQA), finding LID traces a hunchback shape across layers — rising, peaking mid-network, then falling — that closely tracks (shifted one to two layers behind) the layer-wise truthfulness-detection AUROC, peaking at 0.746 and outperforming entropy- and classifier-based baselines; the LID-truthfulness relationship transfers with only modest degradation across QA datasets. See yin-etal-2024-the-local-intrinsic-dimension-of-real-llama-2-activations-traces-a-hunchback-shape-across-layers-that-predicts-generation-truthfulness-on-real-qa-datasets.

Local dimension of fine-tuned representations tracks task success and overfitting onset

Ruppik, von Rohrscheidt, van Niekerk, Heck, Vukovic, Feng, Lin, Lubis, Rieck, Zibrowius & Gasic (2025) apply the TwoNN local-dimension estimator to real RoBERTa-base embeddings, comparing the base masked-LM checkpoint against real task fine-tunes: a TripPy-R dialogue-state tracker fine-tuned on real MultiWOZ 2.1, and a separate fine-tune on real EmoWOZ 7-class emotion recognition. Local dimension drops markedly for the MultiWOZ fine-tune relative to the base model, coinciding (in an auxiliary synthetic grokking task used to validate the estimator) with the onset of rising validation accuracy; on the real EmoWOZ fine-tune, local dimension instead rises starting at the same epoch validation loss begins to increase, tying rising local dimension to overfitting onset rather than task mastery. See ruppik-etal-2025-local-dimension-of-real-fine-tuned-roberta-embeddings-drops-with-successful-dialogue-state-tracking-and-rises-with-overfitting-onset-on-real-emotion-recognition-fine-tuning.

How to detect it

Estimate ID layer-by-layer with a neighbor-distance-ratio estimator (TwoNN or similar), plot it against relative depth, and check for a peak-then-compression shape; cross-check that a semantic-overlap measure (agreement with ground-truth labels among each point’s nearest neighbors) is maximized at or near the ID’s local minimum rather than at the final layer.

Memorized sequences leave a local low-dimensionality signature

Arnold (2025) estimates local intrinsic dimension of hidden-state representations across GPT-Neo (125M/1.3B/2.7B) and BERT while replaying training sequences, finding sequences the model has memorized verbatim show a systematically lower local ID in mid-to-late layers than non-memorized sequences of matched surface statistics, with the gap widening with model scale. A simple ID-based scoring rule detects memorized sequences at accuracy competitive with loss-based membership-inference baselines — extending this node’s central claim (dimensionality profile correlates with what a layer’s representation is “doing”) to a training-data-property axis (memorized vs. not) rather than a depth or training-step axis. See intrinsic-dimension-tracks-memorization-in-language-models.

Effective rank of multi-response, multi-layer embeddings flags hallucination

Wang, Wei, Yue & Sun (2025) directly compute the effective rank (Roy & Vetterli 2007 spectral-entropy dimensionality measure) of a matrix formed by concatenating embeddings sampled across multiple generated responses and multiple layers of Llama-2-7b-chat, Llama-2-13b-chat, and Mistral-7B-v0.1, and use it as a hallucination-detection signal, reaching AUROC around 0.84-0.86 across BioASQ and other QA benchmarks — competitive with or exceeding semantic-entropy and self-consistency baselines. See the-effective-rank-of-multi-response-multi-layer-embedding-matrices-detects-llm-hallucination-at-auroc-0-84-to-0-86-competitive-with-semantic-entropy-baselines.

Combined with anisotropy for hallucination detection

Srey, Wu, Nguyen & Luu (2026) pair a log-pseudo-determinant-of-covariance dimensionality measure with a circular-variance anisotropy measure (see anisotropy) per generated token across Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, and Ministral-8B-Instruct, beating prior uncertainty baselines for hallucination detection across seven benchmarks. See per-token-log-pseudo-determinant-of-cross-layer-covariance-and-circular-variance-of-hidden-states-detect-llm-hallucination-better-than-prior-uncertainty-baselines.

Relative layer-depth of a fixed probe direction shifts with model scale

Manek (2026) extends a diff-of-means evaluation-awareness probe across 11 open-weight models spanning three families (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B), measuring per-layer AUROC on a held-out oversight-detection split rather than a raw dimensionality estimator. The relative depth (layer / total layers) at which decoding AUROC peaks shifts systematically with scale within a family: Qwen2.5 1.5B/3B peak very late (relative depth 0.96-0.97) while 14B/32B peak at the earliest layers (0.021-0.031); Gemma2 shows the same late-to-early shift (0.885 to 0.304) from 2B to 27B; Llama-3.2 stays mid-layer at both sizes tested. This is a different kind of depth-profile claim than the TwoNN/effective-rank trajectories elsewhere in this node — not “how does a layer’s own dimensionality change with depth” but “at what relative depth does a fixed semantic direction become most linearly decodable, and how does that depth itself move as a function of scale” — but it shares the same underlying phenomenon that a network’s geometrically and functionally relevant layer is not fixed at a constant relative position across model sizes. See the-layer-depth-at-which-a-diff-of-means-evaluation-awareness-direction-peaks-in-decoding-accuracy-shifts-from-late-layers-in-small-models-to-the-earliest-layers-in-large-ones-within-the-same-family.

Per-head effective rank tracked across training, rather than across depth

Xu (2026) applies the same effective-rank instinct that motivates this node’s TwoNN/spectral-entropy measures, but to a different axis entirely: rather than tracking overall representation dimensionality across network depth, this work computes the participation ratio (spectral-entropy effective rank) of each individual attention head’s own per-token output activation matrix, integrated across training time. The resulting per-head spectral signal, thresholded via a task-pattern screen and confirmed by causal ablation, identifies which heads are doing specialized circuit computation without any behavioral labels — validated across 7 model configurations spanning a 51M-7B parameter range (dense and MoE architectures, four different pretraining corpora), with the fraction of heads doing identifiable specialized computation conserved at roughly 17-19% across the full scale range, and the signal correctly recovering each of 6 different pretraining seeds’ entirely distinct circuit head-sets without labels. See participation-ratio-spectral-signal-identifies-idiosyncratic-per-seed-attention-head-circuits-without-labels-and-matches-a-conserved-fraction-of-specialized-heads-across-an-8x-scale-range.

A cross-model convergence claim paired with a depth-wise scale trajectory, in a non-language domain

Craig, Selz, Beylich & Tempest (2026) find, alongside a CKA cross-model convergence claim (see Platonic Representation Hypothesis), that GraphCast and Aurora’s processor layers show a consistent depth-wise trajectory: large-spatial-scale changes dominate in early layers, shifting to smaller-spatial-scale changes with increasing depth. Framed via a “particle description” hypothesis (latent variables as particle positions moving under gradient flow toward a learned free-energy minimum), this is a depth-wise structural trajectory documented in AI weather models rather than the language/vision/protein transformers studied elsewhere in this node. See graphcast-and-aurora-show-similar-cka-representational-geometry-with-a-depth-wise-shift-from-large-to-small-spatial-scale-changes.

An information-theoretic effective-features count, cross-checked against prior intrinsic-dimensionality studies

Bereska, Tzifa-Kratira, Samavi & Gavves (2025) define an information- theoretic “effective features” metric F=exp(H)F = \exp(H) (the exponential of the Shannon entropy of SAE feature activation magnitudes) and a superposition ratio ψ=F/N\psi = F/N. Layer-wise effective-feature patterns computed on real pretrained Pythia-70M activations mirror this node’s prior intrinsic-dimensionality findings, and the same metric captures a sharp feature-consolidation transition during grokking. See an-information-theoretic-effective-features-count-from-sae-activation-entropy-tracks-prior-intrinsic-dimensionality-findings-layer-by-layer-in-pythia-70m.

An MST-based estimator, replacing linear-probe accuracy as a representation-quality proxy

Mordacq, Kalogeiton & Oudot (2026) estimate intrinsic dimension of frozen penultimate-layer representations from 33 pretrained SSL vision models — ResNet-50 and ViT-S/B/L/G backbones across joint-embedding (VICReg, DINO), joint-predictive (I-JEPA), combined (iBOT, DINOv2), and vision-language (CLIP, EVA-CLIP) objectives — using a minimum-spanning- tree estimator (dim_MST) rather than the nearest-neighbor-ratio TwoNN estimator used elsewhere in this node. Across ImageNet, iNat-18/21, CIFAR-10/100 and SUN397, dim_MST strongly and consistently correlates with downstream linear-probe accuracy, positioning intrinsic dimension as a cheap proxy for representation quality that doesn’t require training a probe at all. See mst-intrinsic-dimension-of-frozen-ssl-vision-representations-correlates-with-downstream-linear-probe-accuracy-across-33-models.

Local intrinsic dimensionality as a per-input anomaly signature, not a global quality number

Arcos-Holzinger, Erfani, Bailey & Khudanpur (2026) turn intrinsic- dimension estimation into a diagnostic tool: computing per-layer Local Intrinsic Dimensionality (LID) on WavLM and wav2vec 2.0 representations under acoustic perturbation, they find benign low-SNR noise’s LID profile converges back toward the clean-input profile as SNR increases, while adversarial perturbations retain elevated LID in early layers regardless of SNR — a geometric signature distinguishing adversarial from benign inputs (AUROC 0.78-1.00) without needing ground-truth transcripts. See local-intrinsic-dimensionality-of-wavlm-and-wav2vec2-representations-diverges-between-benign-noise-and-adversarial-perturbations-enabling-transcript-free-anomaly-detection.

Prompt-dependent intrinsic dimension in a diffusion model, layer-dependent in its own right

Kvinge, Brown & Godfrey (2023) measure intrinsic dimension of Stable Diffusion’s internal representations at specified bottleneck and latent layers, across denoising steps and varying prompts, finding prompt choice substantially affects the measured dimension. The effect is itself layer-dependent: in certain bottleneck layers, intrinsic dimension correlates with prompt perplexity (via a surrogate language model), while this correlation vanishes in the latent layers — an early (2023), foundational instance of this node’s cross-modality intrinsic-dimension program applied to a generative diffusion model rather than a discriminative or self-supervised one. See stable-diffusions-internal-representation-intrinsic-dimension-depends-on-prompt-and-correlates-with-prompt-perplexity-in-bottleneck-but-not-latent-layers.

Density-peak cluster geometry distinguishes in-context learning from fine-tuning within one LLM

Doimo, Serra, Ansuini & Cazzaniga (2024) combine intrinsic-dimension estimation with Density-peak clustering (Advanced Density Peak) on Llama3-8B’s last-token hidden representations, comparing in-context learning (ICL) and supervised fine-tuning (SFT) layer by layer. Both regimes undergo a sharp two-phase transition around layer 17, marked by a peak in intrinsic dimension and a jump in cluster count/geometry; before the transition, ICL organizes representations into far more (60-70 vs. under 40) and more sharply-separated (core-point fraction ~0.6) semantic clusters than SFT, while after it, SFT instead develops sharper probability modes encoding answer identity — evidence that two training regimes applied to the same base model induce measurably different discovered manifold geometries, not just different downstream accuracy. See icl-and-fine-tuning-in-llama3-8b-undergo-a-shared-layer-17-phase-transition-but-produce-measurably-different-density-peak-cluster-counts-and-separation-before-it.

In-context learning occupies a consistently higher-ID regime than fine-tuning, dissociated from task accuracy

Janapati & Ji (2024) compare in-context learning (ICL), LoRA supervised fine-tuning (SFT), and zero-shot on real Llama-3-8B, Llama-2-13B, Llama-2-7B, and Mistral-7B-v0.3 across 8 tasks, using TwoNN on last-token hidden states. ICL with 5 or more demonstrations induces a consistently higher intrinsic-dimension profile across all layers than SFT or zero-shot, even where SFT achieves higher task accuracy (e.g. MMLU: ICL-10 accuracy 0.531 vs. SFT 0.542) — a scale-and-checkpoint replication of the general finding that different learning paradigms applied to the same base model leave measurably different dimensionality footprints (compare the ICL-vs-SFT cluster-geometry finding above), with ID-vs-demonstration-count itself non-monotonic, rising then plateauing/decreasing past k~5-10. See janapati-ji-2024-in-context-learning-induces-consistently-higher-intrinsic-dimension-than-supervised-fine-tuning-across-real-llama-and-mistral-checkpoints-even-with-lower-task-accuracy.

An early, foundational training-time ID trajectory (expansion then compression)

Razzhigaev, Mikhalchuk, Goncharova, Oseledets, Dimitrov & Kuznetsov (2024) track TwoNN intrinsic dimension across real pretraining checkpoints of Bloom-3B and Pythia-2.8B rather than across depth: “the intrinsic dimension of embeddings increases in the initial phases of training… followed by a compression phase towards the end of training with dimensionality decrease” — an early, foundational instance of the same expansion-then-compression shape this node’s depth-wise entries document, but along the training-time axis (see also the later, more mechanistically detailed Lee et al. 2024 training-dynamics phase-transition entry above). Purely observational. See razzhigaev-etal-2024-embedding-intrinsic-dimension-expands-in-early-pretraining-then-compresses-toward-the-end-tracked-via-twonn-across-bloom-3b-and-pythia-2-8b-checkpoints.

A global-to-local dimension ratio tracks manifold untangling across a reasoning trajectory and model scale

Anderson (2026) pairs PCA global effective dimension (d95d_{95}) with the Levina-Bickel MLE local intrinsic dimension (dmled_{mle}) along chain-of-thought reasoning trajectories in real Llama-3-8B-Instruct and Llama-3.1-70B-Instruct, defining their ratio (Global-to-Local Dimension Ratio, G/L) as an “untangling” statistic distinct from this node’s usual per-layer-depth trajectories: here the axis is generation-time (CoT-step) position, compared across model scale rather than network depth at fixed input. Law-reasoning trajectories show a 10x untangling effect (G/L 9.82x->0.98x) between the two scales. Single-author preprint; observational only. See anderson-2026-a-global-to-local-dimension-ratio-tracks-manifold-untangling-across-cot-reasoning-trajectories-and-model-scale.

Key papers

  • Ansuini, Laio, Macke & Zoccolan (2019). Intrinsic Dimension of Data Representations in Deep Neural Networks. — origin of the TwoNN-based ID-profile analysis, in convolutional (non-transformer) networks.
  • Valeriani, Doimo, Cuturello, Laio, Ansuini & Cazzaniga (2023). The Geometry of Hidden Representations of Large Transformer Models. NeurIPS 2023, arXiv:2302.00294 — extends the analysis to transformers (ESM-2, iGPT), links the ID minimum to peak semantic content, and reports a preliminary Llama-2-70B result.
  • Choi, J., Hwang, G., Cho, H. & Kang, M. (2023). Analyzing the Latent Space of GAN through Local Dimension Estimation. arXiv:2205.13182 — spatially-varying local ID over a real StyleGAN2 latent manifold, correlated with disentanglement.
  • El-Samman, A. M., Husain, I. A., Huynh, M., De Castro, S., Morton, B. & De Baerdemacker, S. (2024). Global geometry of chemical graph neural network representations in terms of chemical moieties. Digital Discovery 3, 544-557 — low effective dimensionality of a real SchNet-family GNN’s QM9 molecular embeddings.
  • Yin, F., Srinivasa, J. & Chang, K.-W. (2024). Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension. ICML 2024, arXiv:2402.18048 — LID hunchback layer profile predicting truthfulness in real Llama-2-7B/13B.
  • Ruppik, B. M., von Rohrscheidt, J., van Niekerk, C., Heck, M., Vukovic, R., Feng, S., Lin, H., Lubis, N., Rieck, B., Zibrowius, M. & Gasic, M. (2025). Less is More: Local Intrinsic Dimensions of Contextual Language Models. NeurIPS 2025, arXiv:2506.01034 — local dimension of fine-tuned RoBERTa tracking task success and overfitting onset.

Found in (36 observations · 40 families)

Earth-Observation / Remote-Sensing Foundation Model

Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning (2026)measured

AlphaEarth satellite embeddings occupy a low-dimensional, locally rotating manifold

Details

Google AlphaEarth's 64-dimensional satellite embeddings have a participation-ratio effective dimensionality of 13.3 and a local intrinsic dimensionality of about 10 across ~12.1M CONUS samples (2017-2023) [rahman-etal-2026-characterizing-alphaearth-embedding-geometry] This intrinsic-to-ambient ratio is higher than comparable geographic implicit neural representations, whose prior ID was 2-10 in 256-512 ambient dimensions [rahman-etal-2026-characterizing-alphaearth-embedding-geometry] Tangent spaces between adjacent probe locations rotate more than 60 degrees at 84% of sampled locations, indicating a strongly curved (non-affine) manifold [rahman-etal-2026-characterizing-alphaearth-embedding-geometry] A separate local-versus-global comparison finds a mean |cos theta| of 0.169 between local tangent spaces and the global principal axes, a distinct metric from the adjacent-probe rotation [rahman-etal-2026-characterizing-alphaearth-embedding-geometry] Tangent-space instability is highest in mountainous regions with steep environmental gradients [rahman-etal-2026-characterizing-alphaearth-embedding-geometry]

models: AlphaEarth (satellite/Earth-observation foundation model) · method: PCA, Intrinsic dimension estimation (TwoNN)

Pythia

Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability (2025)measured

SAE activation-entropy effective-features count mirrors intrinsic-dimensionality studies

Details

Bereska et al. define effective-features F = exp(Shannon entropy of SAE feature-activation magnitudes) and a superposition ratio psi = F/N as an effective-dimensionality measure of a representation [bereska-etal-2025-superposition-as-lossy-compression] Computed layer-by-layer on real pretrained Pythia-70M activations, the effective-features profile mirrors prior intrinsic-dimensionality studies of the same model [bereska-etal-2025-superposition-as-lossy-compression] The same metric captures a sharp feature-consolidation transition during grokking [bereska-etal-2025-superposition-as-lossy-compression] Dropout systematically reduces the number of effective features [bereska-etal-2025-superposition-as-lossy-compression]

models: Pythia-70M · method: Sparse Autoencoders (SAE)
Representational Curvature Modulates Behavioral Uncertainty in Large Language Models (2026)measured

Trajectory-subspace curvature perturbations causally modulate next-token entropy

Details

King et al. define contextual curvature as a 3-token backward average of the angle between consecutive residual-stream displacement vectors, extending the Hosseini-Fedorenko straightening measure [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Across GPT-2 XL and Pythia-2.8B, contextual curvature predicts next-token entropy (Pearson r peaking ~0.15), rising across layers to a peak near the middle-layer curvature minimum [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] In Pythia's training the coupling is absent at 0-0.07% of a 300B-token run and emerges sharply around 0.7% [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Only perturbations restricted to the recent-displacement subspace or its 2D plane causally move entropy; random, random-subspace, activation-PCA, and full-space perturbations of matched magnitude do not [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Training a 12-layer model from scratch with a curvature-regularizing auxiliary loss lowers token-level entropy without degrading validation loss [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty]

models: Pythia-2.8B · method: Geometric analysis, Causal interventions (steering)
Abstraction Induces the Brain Alignment of Language and Speech Models (2026)measured

Layer intrinsic-dimension peak aligns with the best brain-predicting layer

Details

Cheng et al. estimate per-layer intrinsic dimension via the GRIDE estimator in OPT (125M/1.3B/13B), Pythia (160M/410M/6.9B), WavLM (base-plus, large), and Whisper (large encoder) [cheng-etal-2026-abstraction-brain-alignment] Layerwise intrinsic dimension correlates with how well each layer predicts real human brain responses (fMRI rho=0.76, ECoG rho=0.43, both p<0.05) [cheng-etal-2026-abstraction-brain-alignment] The ID-peak layer and the best-brain-predicting layer are usually within 0-1 layers of each other [cheng-etal-2026-abstraction-brain-alignment] Brain-tuning WavLM's best layer (layer 9) to predict fMRI causally raises both its intrinsic dimension and its semantic content, while a random-Fourier-features control shows raised ID alone is insufficient [cheng-etal-2026-abstraction-brain-alignment]

models: Pythia-160M, Pythia-410M, Pythia-6.9B · method: Intrinsic dimension estimation (TwoNN), Causal interventions (steering)
Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024)measured

Nonlinear intrinsic dimension phase-transitions with emergent zero-shot competence

Details

Lee et al. track per-layer nonlinear intrinsic dimension (TwoNN) and linear PCA effective dimension across pretraining of Pythia-410M/1.4B/6.9B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] On synthetic data both measures scale with dataset compositionality in fully-trained models, but nonlinear ID shows a sharp phase transition around step ~10^3 coinciding with the onset of zero-shot competence [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] PCA effective dimension tracks superficial/Kolmogorov complexity (gzip compressibility, early-training Spearman rho near 1.0) while TwoNN ID never correlates with gzip [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] The linear-superficial versus nonlinear-semantic dissociation generalizes to fully-trained Llama-3-8B and Mistral-7B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime]

models: Pythia-410M, Pythia-1.4B, Pythia-6.9B · method: Intrinsic dimension estimation (TwoNN), PCA
Rethinking Intrinsic Dimension Estimation in Neural Representations (2026)measured

The hump-shaped ID profile is likely a TwoNN estimator artifact

Details

Schulte & Rugamer prove that standard network layers (linear/conv, ReLU, softmax, pooling, residual, normalization, self-attention) are Lipschitz maps, so true pointwise and Hausdorff intrinsic dimension can only stay equal or decrease across a layer, never increase [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Yet the standard TwoNN/MLE/GRIDE pipeline on ResNet-34, Llama-3.1-8B, Mistral-7B-v0.3, and Pythia-6.9B reproduces the familiar hump-shaped expansion-then-compression ID profile, a pattern their theorem proves cannot be the true ID [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Investigating what the estimators actually respond to (neighbor distances, ambient dimension, cosine similarity, norm, entropy), they conclude the mid-layer-ID-peak abstraction narrative is very likely a systematic estimator artifact [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] The caution applies specifically to TwoNN and related nearest-neighbor-ratio estimators, not to MST-based, spectral-entropy/effective-rank, or SVD-based measures [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation]

models: Pythia-6.9B · method: Intrinsic dimension estimation (TwoNN)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026)measured

A participation-ratio spectral signal finds attention circuits label-free across scale

Details

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

models: Pythia-160M, Pythia-410M, Pythia-1B · method: Participation-ratio spectral signal (per-head effective rank over training), Causal interventions (steering)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024)measured

Embedding intrinsic dimension expands early in pretraining, then compresses

Details

Razzhigaev et al. apply TwoNN (cross-validated against Manifold-Adaptive Dimension Estimation and the Method of Moments) to embeddings sampled across real pretraining checkpoints of Bloom-3B and Pythia-2.8B [razzhigaev-etal-2024-shape-of-learning] Intrinsic dimension rises in the initial phase of training, then compresses toward the end, a training-time (not depth-wise) dimensionality trajectory [razzhigaev-etal-2024-shape-of-learning] The analysis is purely observational, with no causal intervention [razzhigaev-etal-2024-shape-of-learning]

models: Pythia-2.8B · method: Intrinsic dimension estimation (TwoNN)

Llama

The Geometry of Thought: How Scale Restructures Reasoning in Large Language Models (2026)measured

A global-to-local PCA-dimension ratio tracks reasoning-trajectory untangling

Details

Anderson defines a Global-to-Local Dimension Ratio as PCA d95 measured globally over a chain-of-thought trajectory divided by PCA d95 measured locally, with the Levina-Bickel MLE intrinsic dimension computed separately as a check [anderson-2026-geometry-of-thought] A high ratio means the trajectory fills a large ambient subspace while staying locally low-dimensional and folded; a ratio near 1 means an unfolded, roughly linear trajectory [anderson-2026-geometry-of-thought] Law-reasoning trajectories untangle roughly 10-fold, the ratio falling from 9.82 to 0.98 between Llama-3-8B-Instruct and Llama-3.1-70B-Instruct [anderson-2026-geometry-of-thought] The dimension profile is tracked along the generation-time (CoT-step) axis and compared across model scale, rather than across network depth [anderson-2026-geometry-of-thought] No causal intervention is performed [anderson-2026-geometry-of-thought]

models: Llama-3-8B-Instruct, Llama-3.1-70B-Instruct · method: Intrinsic dimension estimation (TwoNN), PCA
The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language Models (2024)measured

ICL and fine-tuning share a layer-17 transition but differ before it

Details

Doimo et al. apply density-peak clustering and intrinsic-dimension estimation to Llama3-8B last-token representations, comparing in-context learning and supervised fine-tuning layer by layer [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] Both regimes undergo a sharp two-phase transition around layer 17, marked by an ID peak and a jump in density-peak cluster geometry [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] Before the transition, ICL organizes representations into far more clusters (60-70 versus under 40) and more sharply separated ones (core-point fraction ~0.6) than SFT [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] After the transition, SFT develops sharper probability modes encoding answer identity, showing the two regimes induce measurably different cluster geometry within the same base model [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning]

models: Llama-3-8B · method: Density-peak clustering (Advanced Density Peak), Intrinsic dimension estimation (TwoNN)
The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024)measured

Intrinsic dimension expands then compresses; semantics peak at the minimum

Details

Valeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]

models: Llama-2-70B · method: Intrinsic dimension estimation (TwoNN), Neighborhood overlap, PCA, Sparse Autoencoders (SAE)
A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension (2024)measured

In-context learning induces higher intrinsic dimension than fine-tuning

Details

Janapati & Ji apply TwoNN to last-token representations of Llama-3-8B, Llama-2-13B, Llama-2-7B, and Mistral-7B-v0.3 across 8 tasks, comparing ICL, LoRA fine-tuning, and zero-shot [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] ICL with k>=5 demonstrations induces a consistently higher intrinsic-dimension profile across all layers than SFT or zero-shot [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] This holds even though SFT reaches higher task accuracy (e.g. MMLU 0.542 vs ICL 0.531), dissociating performance from manifold dimensionality [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] ID versus number of demonstrations is non-monotonic, rising then plateauing past k approximately 5-10 [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension]

models: Llama-3-8B, Llama-2-13B, Llama-2-7B · method: Intrinsic dimension estimation (TwoNN)
Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024)measured

Nonlinear intrinsic dimension phase-transitions with emergent zero-shot competence

Details

Lee et al. track per-layer nonlinear intrinsic dimension (TwoNN) and linear PCA effective dimension across pretraining of Pythia-410M/1.4B/6.9B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] On synthetic data both measures scale with dataset compositionality in fully-trained models, but nonlinear ID shows a sharp phase transition around step ~10^3 coinciding with the onset of zero-shot competence [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] PCA effective dimension tracks superficial/Kolmogorov complexity (gzip compressibility, early-training Spearman rho near 1.0) while TwoNN ID never correlates with gzip [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] The linear-superficial versus nonlinear-semantic dissociation generalizes to fully-trained Llama-3-8B and Mistral-7B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime]

models: Llama-3-8B · method: Intrinsic dimension estimation (TwoNN), PCA
Shared Global and Local Geometry of Language Model Embeddings (2025)measured

Token-embedding orientation is shared within families but drops across them

Details

Lee et al. measure token-embedding geometry across the GPT-2, Llama-3, and Gemma-2 model families [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] Relative orientation/cosine structure is near-identical within a model family but drops sharply across families trained on different data [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] A k-NN-neighborhood PCA intrinsic-dimension estimator finds low-ID tokens form semantically coherent clusters while high-ID tokens do not [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] EMB2EMB, a linear least-squares (OLS) map fit over 100k shared tokens, transfers CAA-style steering vectors (refusal, sycophancy, corrigibility) between differently-sized models [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings]

models: Llama-3.1-8B-Instruct · method: PCA, Intrinsic dimension estimation (TwoNN), Cross-model direction transfer via ridge regression
Rethinking Intrinsic Dimension Estimation in Neural Representations (2026)measured

The hump-shaped ID profile is likely a TwoNN estimator artifact

Details

Schulte & Rugamer prove that standard network layers (linear/conv, ReLU, softmax, pooling, residual, normalization, self-attention) are Lipschitz maps, so true pointwise and Hausdorff intrinsic dimension can only stay equal or decrease across a layer, never increase [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Yet the standard TwoNN/MLE/GRIDE pipeline on ResNet-34, Llama-3.1-8B, Mistral-7B-v0.3, and Pythia-6.9B reproduces the familiar hump-shaped expansion-then-compression ID profile, a pattern their theorem proves cannot be the true ID [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Investigating what the estimators actually respond to (neighbor distances, ambient dimension, cosine similarity, norm, entropy), they conclude the mid-layer-ID-peak abstraction narrative is very likely a systematic estimator artifact [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] The caution applies specifically to TwoNN and related nearest-neighbor-ratio estimators, not to MST-based, spectral-entropy/effective-rank, or SVD-based measures [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation]

models: Llama-3.1-8B · method: Intrinsic dimension estimation (TwoNN)
The Geometry of Reasoning: Flowing Logics in Representation Space (2026)measured

Reasoning-step curvature tracks logical structure, not surface form

Details

Zhou et al. model each reasoning step as a mean-pooled final-layer hidden state and compute the Menger curvature of the step sequence across 2,430 sequences that vary logic, topic, and language independently [zhou-etal-2026-geometry-of-reasoning] Curvature similarity is consistently higher for logic-matched pairs than topic- or language-matched pairs (e.g. Qwen3-0.6B: logic 0.53 vs topic 0.11 vs language 0.13; Llama3-8B: 0.58 vs 0.13 vs 0.17) [zhou-etal-2026-geometry-of-reasoning] Raw position similarity is instead dominated by language (Qwen3-0.6B: language 0.85 vs logic 0.26), so second-order curvature, not position, is organized by logical structure [zhou-etal-2026-geometry-of-reasoning] A step-shuffle control collapses velocity and curvature similarity toward zero while leaving position similarity high, and no causal intervention is performed [zhou-etal-2026-geometry-of-reasoning]

models: Llama-3-8B · method: Geometric analysis, PCA
Revisiting Hallucination Detection with Effective Rank-based Uncertainty (2025)measured

Multi-response effective rank detects hallucination at AUROC 0.84-0.86

Details

Wang et al. compute the effective rank (Roy-Vetterli spectral-entropy dimensionality) of a matrix of embeddings sampled across multiple generated responses and layers of Llama-2-7b-chat, Llama-2-13b-chat, and Mistral-7B-v0.1 [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Used as a hallucination signal it reaches AUROC around 0.84-0.86 across QA benchmarks, competitive with or exceeding semantic-entropy and self-consistency baselines [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] The 13B BioASQ configuration scores 0.8234, slightly below the headline 0.84-0.86 range [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Ablations vary the number of generations and the layer-selection strategy (middle layer vs last-5 layers) [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty]

models: Llama-2-7B-Chat, Llama-2-13B-Chat · method: Intrinsic dimension estimation (TwoNN)
Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs (2026)measured

Evaluation-awareness probe depth shifts late-to-early with model scale

Details

Manek projects a diff-of-means evaluation-awareness direction (from 203 contrastive prompt pairs) onto residual-stream activations at every layer across 11 open-weight models (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] The relative depth at which per-layer decoding AUROC peaks shifts from late layers in small models to the earliest layers in large ones (Qwen2.5 1.5B/3B peak at 0.96-0.97 vs 14B/32B at 0.021-0.031) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Gemma2 shows the same late-to-early shift (0.885 to 0.304) while Llama-3.2 stays mid-layer across both tested sizes [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Peak AUROC ranges 0.586-0.873 and is non-monotonic with scale within a family [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale]

models: Llama-3.2-1B, Llama-3.2-3B · method: Direction Extraction, Linear probing, Geometric analysis
Lines of Thought in Large Language Models (2024)measured

Token trajectories trace a curved low-dimensional Langevin manifold

Details

Sarfati et al. track each token's hidden-state trajectory through GPT-2-Medium's 24 layers (replicated on Llama-2-7B, Mistral-7B-v0.1, Llama-3.2-1B/3B) and find ensembles cluster on a low-dimensional curved manifold [sarfati-etal-2024-lines-of-thought] Truncating to the top ~256 of 1024 per-layer SVD dimensions preserves ~90% of the next-token distribution's information [sarfati-etal-2024-lines-of-thought] Layer-to-layer evolution follows a rotation-and-stretch law plus exponentially-growing Gaussian noise, a linear Langevin/Ornstein-Uhlenbeck SDE extracted from ensemble statistics [sarfati-etal-2024-lines-of-thought] Simulated SDE trajectories reproduce real statistics (a linear SVM separates them at near-chance 46-61%), while untrained-model trajectories move in straight parallel lines, so the structure is training-dependent [sarfati-etal-2024-lines-of-thought]

models: Llama-2-7B, Llama-3.2-1B, Llama-3.2-3B · method: SVD, Geometric analysis
Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension (2024)measured

The local intrinsic dimension of real Llama-2 activations traces a hunchback shape across layers that predicts generation truthfulness on real QA datasets

Details

Local Intrinsic Dimension (LID, via the GeoMLE estimator) is computed on the per-layer hidden activations of real Llama-2-7B and Llama-2-13B, generating answers zero-shot and few-shot on four real QA datasets -- TriviaQA, HotpotQA, TydiQA-GP, and CoQA [yin-etal-2024-truthfulness-local-intrinsic-dimension] Aggregated across Llama-2-7B's 30 layers, LID follows a hunchback shape -- rising in early layers, peaking mid-network, then declining -- that closely tracks (shifted by one or two layers behind) the layer-wise truthfulness-detection AUROC, which peaks at 0.746 averaged across datasets, outperforming entropy- and classifier-based baselines [yin-etal-2024-truthfulness-local-intrinsic-dimension] A truthfulness detector trained on one QA dataset's LID profile transfers with only modest degradation to a different QA dataset (e.g. CoQA AUROC 0.763 with same-dataset neighbors versus 0.747 with TriviaQA neighbors), indicating the LID-truthfulness relationship is not dataset-specific [yin-etal-2024-truthfulness-local-intrinsic-dimension]

models: Llama-2-7B, Llama-2-13B · method: Local Intrinsic Dimensionality (LID) under perturbation

AlexNet

Intrinsic Dimension of Data Representations in Deep Neural Networks (2019)measured

CNN intrinsic dimension is hunchback-shaped; last-layer ID predicts accuracy

Details

Ansuini et al. estimate TwoNN intrinsic dimension of layer-wise activations in real trained ImageNet CNNs (AlexNet, VGG, ResNet), finding ID rises sharply in early layers then contracts to a low plateau, a hunchback profile [ansuini-etal-2019-intrinsic-dimension-of-data-representations] Intrinsic dimension is orders of magnitude below the nominal unit count at every layer [ansuini-etal-2019-intrinsic-dimension-of-data-representations] The last hidden layer's intrinsic dimension predicts the network's own test accuracy across architectures (r approximately 0.94) [ansuini-etal-2019-intrinsic-dimension-of-data-representations] This foundational result predates and matches the expansion-then-compression pattern later confirmed in Transformers [ansuini-etal-2019-intrinsic-dimension-of-data-representations] No causal intervention is performed [ansuini-etal-2019-intrinsic-dimension-of-data-representations]

models: AlexNet (ImageNet image classifier, supervised) · method: Intrinsic dimension estimation (TwoNN)

VGG

Intrinsic Dimension of Data Representations in Deep Neural Networks (2019)measured

CNN intrinsic dimension is hunchback-shaped; last-layer ID predicts accuracy

Details

Ansuini et al. estimate TwoNN intrinsic dimension of layer-wise activations in real trained ImageNet CNNs (AlexNet, VGG, ResNet), finding ID rises sharply in early layers then contracts to a low plateau, a hunchback profile [ansuini-etal-2019-intrinsic-dimension-of-data-representations] Intrinsic dimension is orders of magnitude below the nominal unit count at every layer [ansuini-etal-2019-intrinsic-dimension-of-data-representations] The last hidden layer's intrinsic dimension predicts the network's own test accuracy across architectures (r approximately 0.94) [ansuini-etal-2019-intrinsic-dimension-of-data-representations] This foundational result predates and matches the expansion-then-compression pattern later confirmed in Transformers [ansuini-etal-2019-intrinsic-dimension-of-data-representations] No causal intervention is performed [ansuini-etal-2019-intrinsic-dimension-of-data-representations]

models: VGG (image classifier, various depths) · method: Intrinsic dimension estimation (TwoNN)

ResNet

Intrinsic Dimension of Data Representations in Deep Neural Networks (2019)measured

CNN intrinsic dimension is hunchback-shaped; last-layer ID predicts accuracy

Details

Ansuini et al. estimate TwoNN intrinsic dimension of layer-wise activations in real trained ImageNet CNNs (AlexNet, VGG, ResNet), finding ID rises sharply in early layers then contracts to a low plateau, a hunchback profile [ansuini-etal-2019-intrinsic-dimension-of-data-representations] Intrinsic dimension is orders of magnitude below the nominal unit count at every layer [ansuini-etal-2019-intrinsic-dimension-of-data-representations] The last hidden layer's intrinsic dimension predicts the network's own test accuracy across architectures (r approximately 0.94) [ansuini-etal-2019-intrinsic-dimension-of-data-representations] This foundational result predates and matches the expansion-then-compression pattern later confirmed in Transformers [ansuini-etal-2019-intrinsic-dimension-of-data-representations] No causal intervention is performed [ansuini-etal-2019-intrinsic-dimension-of-data-representations]

models: ResNet (image classifier, various depths) · method: Intrinsic dimension estimation (TwoNN)
Rethinking Intrinsic Dimension Estimation in Neural Representations (2026)measured

The hump-shaped ID profile is likely a TwoNN estimator artifact

Details

Schulte & Rugamer prove that standard network layers (linear/conv, ReLU, softmax, pooling, residual, normalization, self-attention) are Lipschitz maps, so true pointwise and Hausdorff intrinsic dimension can only stay equal or decrease across a layer, never increase [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Yet the standard TwoNN/MLE/GRIDE pipeline on ResNet-34, Llama-3.1-8B, Mistral-7B-v0.3, and Pythia-6.9B reproduces the familiar hump-shaped expansion-then-compression ID profile, a pattern their theorem proves cannot be the true ID [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Investigating what the estimators actually respond to (neighbor distances, ambient dimension, cosine similarity, norm, entropy), they conclude the mid-layer-ID-peak abstraction narrative is very likely a systematic estimator artifact [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] The caution applies specifically to TwoNN and related nearest-neighbor-ratio estimators, not to MST-based, spectral-entropy/effective-rank, or SVD-based measures [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation]

models: ResNet-34 (supervised, ImageNet) · method: Intrinsic dimension estimation (TwoNN)
Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and the Human Brain (2026)measured

Lower local intrinsic dimension predicts vision-model alignment and generalization

Details

Yu et al. estimate local intrinsic dimension via the Levina-Bickel maximum-likelihood kNN estimator across a pool of 91 vision models (ConvNeXt, ResNet, ResMLP, ViT), with a 51-model architecture-balanced subset for cross-architecture comparison [yu-etal-2026-local-intrinsic-dimension-alignment] Local ID is significantly negatively correlated with AI-AI representational alignment, AI-brain alignment (fMRI, Natural Scenes Dataset), and ImageNet-1K generalization [yu-etal-2026-local-intrinsic-dimension-alignment] The correlation is strongest at small (local) neighborhood scale K and weakens toward global estimates, with a matched-subsample control confirming genuine local structure [yu-etal-2026-local-intrinsic-dimension-alignment] Increasing model capacity and training-data scale systematically reduces local ID (local ID vs log-parameters Corr=-0.756, p<0.001) [yu-etal-2026-local-intrinsic-dimension-alignment] PCA is used only as a top-300-component control to equalize ambient dimensionality, not as the ID estimator, and robustness is checked with the MOM and MADA estimators [yu-etal-2026-local-intrinsic-dimension-alignment]

models: ResNet (image classifier, various depths) · method: Local Intrinsic Dimensionality (LID) under perturbation, PCA

ESM-2

Towards Understanding the Shape of Representations in Protein Language Models (2026)measured

Protein-LM shape spaces expand then compress at low absolute dimension

Details

Beshkov & Malthe-Sorenssen treat each protein's per-residue PLM representation as a curve, map it to a square-root-velocity shape space to quotient out rotation and translation, and measure Frechet radius and tangent-PCA effective dimension layer-by-layer across ESM2 (35M/150M/650M/3B) and Ankh on 1,377 SCOPe proteins [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] Frechet radius decreases with depth and is far smaller for PLM shape space than for real 3D protein structure, largely independent of model size [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] Tangent-PCA effective dimension shows an expansion-then-compression trajectory across layers, more pronounced for larger models, at far lower absolute dimension than PCA on flattened pointwise activations [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] A graph-filtration comparison finds 3D structure is most faithfully encoded at short (2-residue) and moderate (8-residue) context, peaking just before the final layer in every model [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models]

models: ESM-2 (35M), ESM-2 (150M), ESM-2 (650M), ESM-2 (3B) · method: SRV shape-space Fréchet radius and tangent-PCA effective dimension
The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024)measured

Intrinsic dimension expands then compresses; semantics peak at the minimum

Details

Valeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]

models: ESM-2 (35M), ESM-2 (650M), ESM-2 (3B) · method: Intrinsic dimension estimation (TwoNN), Neighborhood overlap, PCA, Sparse Autoencoders (SAE)

Ankh

Towards Understanding the Shape of Representations in Protein Language Models (2026)measured

Protein-LM shape spaces expand then compress at low absolute dimension

Details

Beshkov & Malthe-Sorenssen treat each protein's per-residue PLM representation as a curve, map it to a square-root-velocity shape space to quotient out rotation and translation, and measure Frechet radius and tangent-PCA effective dimension layer-by-layer across ESM2 (35M/150M/650M/3B) and Ankh on 1,377 SCOPe proteins [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] Frechet radius decreases with depth and is far smaller for PLM shape space than for real 3D protein structure, largely independent of model size [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] Tangent-PCA effective dimension shows an expansion-then-compression trajectory across layers, more pronounced for larger models, at far lower absolute dimension than PCA on flattened pointwise activations [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models] A graph-filtration comparison finds 3D structure is most faithfully encoded at short (2-residue) and moderate (8-residue) context, peaking just before the final layer in every model [beshkov-malthe-sorenssen-2026-towards-understanding-the-shape-of-representations-in-protein-language-models]

models: Ankh (base) · method: SRV shape-space Fréchet radius and tangent-PCA effective dimension

StyleGAN

Analyzing the Latent Space of GAN through Local Dimension Estimation (2023)measured

Local intrinsic dimension of a real StyleGAN2 latent manifold varies across latent space and correlates with disentanglement

Details

Treating a real pretrained StyleGAN/StyleGAN2 generator's intermediate latent space as a Riemannian manifold, local intrinsic dimension is estimated at each sampled latent point via the generator's own local tangent-space structure [choi-etal-2023-local-intrinsic-dimension-gan-latent] Local intrinsic dimension varies substantially across the latent manifold rather than being constant, and is interpreted as the number of independent semantic variations available at that latent point [choi-etal-2023-local-intrinsic-dimension-gan-latent] A derived tangent-space-inconsistency metric (Distortion) correlates highly with supervised disentanglement scores despite requiring no attribute labels [choi-etal-2023-local-intrinsic-dimension-gan-latent]

models: StyleGAN2 (trained on FFHQ, 1024x1024) · method: Generator/decoder Jacobian pullback metric

Gemma

Context Structure Reshapes the Representational Geometry of Language Models (2026)measured

Context straightens representational trajectories, but only in continual prediction

Details

Hosseini et al. study Gemma-2-27B residual-stream trajectories across four task families using trajectory curvature (the angle between consecutive per-token transition vectors), participation-ratio effective dimensionality, and PC1/PC2 elongation [hosseini-etal-2026-context-structure-geometry] Trajectories straighten (curvature falls) as context grows in continual-prediction settings, in natural language and grid-world tasks (long- vs short-context t=-7.46, d=0.75; latent grid t=-11.62, d=1.16) [hosseini-etal-2026-context-structure-geometry] Straightening is inconsistent or absent in structured-prediction tasks: few-shot learning straightens only in the transition phase (F=152.7) and not during answer generation (F=0.67, p=0.72) [hosseini-etal-2026-context-structure-geometry] Straightening correlates with behavioral output-logit gaps in grid-world tasks (r=0.99), but the paper performs no causal intervention [hosseini-etal-2026-context-structure-geometry]

models: Gemma-2-27B · method: PCA, Geometric analysis
The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws (2026)measured

Manifold curvature predicts sparse-autoencoder reconstruction-loss scaling floors

Details

Zaher et al. estimate per-layer intrinsic dimension (TwoNN) and multi-scale curvature (local-PCA tangent-variation and heterogeneity) of activation manifolds in Gemma-2 2B and 9B [zaher-etal-2026-geometric-wall] Regressed against fitted SAE scaling-law parameters from 844 Gemma Scope JumpReLU checkpoints, multi-scale curvature kappa_ms is the single dominant predictor of the scaling-law exponent, more than intrinsic dimension alone [zaher-etal-2026-geometric-wall] The full geometric model reaches leave-one-out R^2=0.869 (9B) and 0.976 (2B), and features fit on one model predict the other's scaling law at cross-model R^2>0.92 [zaher-etal-2026-geometric-wall] The analysis is correlational, with no causal manifold-deformation intervention [zaher-etal-2026-geometric-wall]

models: Gemma-2-2B, Gemma-2-9B · method: Intrinsic dimension estimation (TwoNN), Geometric analysis, Sparse Autoencoders (SAE)
The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024)measured

Intrinsic dimension expands then compresses; semantics peak at the minimum

Details

Valeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]

models: Gemma-2-2B · method: Intrinsic dimension estimation (TwoNN), Neighborhood overlap, PCA, Sparse Autoencoders (SAE)
Shared Global and Local Geometry of Language Model Embeddings (2025)measured

Token-embedding orientation is shared within families but drops across them

Details

Lee et al. measure token-embedding geometry across the GPT-2, Llama-3, and Gemma-2 model families [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] Relative orientation/cosine structure is near-identical within a model family but drops sharply across families trained on different data [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] A k-NN-neighborhood PCA intrinsic-dimension estimator finds low-ID tokens form semantically coherent clusters while high-ID tokens do not [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] EMB2EMB, a linear least-squares (OLS) map fit over 100k shared tokens, transfers CAA-style steering vectors (refusal, sycophancy, corrigibility) between differently-sized models [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings]

models: Gemma-2-9B-it · method: PCA, Intrinsic dimension estimation (TwoNN), Cross-model direction transfer via ridge regression
Reasoning emerges from constrained inference manifolds in large language models (2026)measured

Chain-of-thought trajectories collapse to under ten intrinsic dimensions

Details

Ma et al. estimate intrinsic dimension with the TLE estimator (k=20) for the static vocabulary-embedding cloud (D_world) and for autoregressive last-token chain-of-thought trajectories (D_stim) across Qwen2.5, Qwen3, Gemma 3, and DeepSeek-R1-Distill-Qwen families [ma-etal-2026-constrained-inference-manifolds] Deep-layer reasoning trajectories collapse onto a manifold of fewer than ten intrinsic degrees of freedom despite thousands of ambient dimensions, while the vocabulary cloud stays near ambient dimension [ma-etal-2026-constrained-inference-manifolds] A composite reasoning-health diagnostic H = log(D_world)*V/exp(0.1*D_stim), combining an information-volume term V, correlates with benchmark scores at Spearman rho>0.9 [ma-etal-2026-constrained-inference-manifolds] Token-shuffle, non-cognitive-prompt, and truncation controls rule out trivial confounds, and no causal intervention is performed [ma-etal-2026-constrained-inference-manifolds]

models: Gemma 3 1B, Gemma 3 4B, Gemma 3 12B, Gemma 3 27B · method: Intrinsic dimension estimation (TwoNN), Geometric analysis
Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs (2026)measured

Evaluation-awareness probe depth shifts late-to-early with model scale

Details

Manek projects a diff-of-means evaluation-awareness direction (from 203 contrastive prompt pairs) onto residual-stream activations at every layer across 11 open-weight models (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] The relative depth at which per-layer decoding AUROC peaks shifts from late layers in small models to the earliest layers in large ones (Qwen2.5 1.5B/3B peak at 0.96-0.97 vs 14B/32B at 0.021-0.031) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Gemma2 shows the same late-to-early shift (0.885 to 0.304) while Llama-3.2 stays mid-layer across both tested sizes [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Peak AUROC ranges 0.586-0.873 and is non-monotonic with scale within a family [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale]

models: Gemma-2-2B, Gemma-2-9B, Gemma-2-27B · method: Direction Extraction, Linear probing, Geometric analysis

SchNet (continuous-filter convolutional GNN for molecules)

Global geometry of chemical graph neural network representations in terms of chemical moieties (2024)measured

A real SchNet-family GNN's 128-parameter QM9 molecular embedding space reduces to around 5 effective parameters, with sharp linear boundaries separating chemical moieties

Details

Dimension reduction and linear discriminant analysis applied to the hidden-layer embedding vectors of a real SchNet-family GNN trained on the real QM9 molecular dataset show the fully-trained 128-parameter embedding space reduces to a low parametric space of around 5 important parameters [el-samman-etal-2024-global-geometry-chemical-gnn-moieties] Sharp linear boundaries separate chemical moieties within this reduced embedding space, with classification error below 5x10^-4 [el-samman-etal-2024-global-geometry-chemical-gnn-moieties] Euclidean distance in the embedding space functions as a molecular similarity measure competitive with the hand-engineered SOAP descriptor, and embedding coordinates linearly predict pKa and NMR chemical-shift observables [el-samman-etal-2024-global-geometry-chemical-gnn-moieties]

models: SchNet-family GNN, trained on QM9 (El-Samman, Husain, Huynh, De Castro, Morton & De Baerdemacker) · method:

Stable Diffusion

ELROND: Exploring and Decomposing Intrinsic Capabilities of Diffusion Models (2026)measured

Concept-direction local ID tracks generality and restores distilled-diffusion diversity

Details

Skiers et al. backpropagate differences between stochastic image realizations of the same prompt in real SDXL, then decompose the gradient directions via PCA or a sparse autoencoder into concept-specific latent directions [skiers-etal-2026-elrond-diffusion-concept-decomposition] The local intrinsic dimension of each concept's manifold tracks concept generality: general concepts (e.g. "Dog") show higher LID than their hyponyms (e.g. "Poodle"), validated against WordNet pairs [skiers-etal-2026-elrond-diffusion-concept-decomposition] Adding the discovered directions into the distilled, mode-collapsed SDXL-DMD student causally steers single concepts and, combined, restores output diversity toward the teacher (FID improves), most when directions come from the teacher [skiers-etal-2026-elrond-diffusion-concept-decomposition] Equal-norm random directions are semantically inert by comparison [skiers-etal-2026-elrond-diffusion-concept-decomposition]

models: Stable Diffusion XL, SDXL-DMD (4-step distilled) · method: Local Intrinsic Dimensionality (LID) under perturbation, PCA, Sparse Autoencoders (SAE), Causal interventions (steering)
Exploring the Representation Manifolds of Stable Diffusion Through the Lens of Intrinsic Dimension (2023)measured

Stable Diffusion's intrinsic dimension tracks prompt perplexity in bottleneck layers

Details

Kvinge et al. estimate the intrinsic dimension of Stable Diffusion's internal representations at bottleneck and latent layers across denoising steps and prompts, finding prompt choice substantially shifts the measured ID [kvinge-etal-2023-exploring-representation-manifolds-of-stable-diffusion] In certain bottleneck layers ID correlates with prompt perplexity (via a surrogate language model), but this correlation vanishes in the latent layers [kvinge-etal-2023-exploring-representation-manifolds-of-stable-diffusion] The representation manifold thus responds to linguistic complexity differently depending on where along the architecture it is measured [kvinge-etal-2023-exploring-representation-manifolds-of-stable-diffusion]

models: Stable Diffusion v1.4 · method: Intrinsic dimension estimation (TwoNN)

Weather/Climate Foundation Model

The physics of AI weather models (2026)measured

GraphCast and Aurora share CKA geometry with a depth-wise scale shift

Details

Craig et al. compute Centered Kernel Alignment between AI weather models and find forecast skill correlates with cross-model representational alignment [craig-etal-2026-the-physics-of-ai-weather-models] GraphCast and Aurora represent the atmosphere similarly despite differing architectures and capacity [craig-etal-2026-the-physics-of-ai-weather-models] They propose a "particle description" in which latent variables move under gradient flow toward a minimum of a learned free-energy functional [craig-etal-2026-the-physics-of-ai-weather-models] Consistent with this, processor layers shift from large-spatial-scale changes early to small-spatial-scale changes with depth [craig-etal-2026-the-physics-of-ai-weather-models]

models: GraphCast (weather foundation model), Aurora (weather foundation model) · method: Centered Kernel Alignment (CKA)

Image GPT

The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024)measured

Intrinsic dimension expands then compresses; semantics peak at the minimum

Details

Valeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]

models: iGPT-S, iGPT-M, iGPT-L · method: Intrinsic dimension estimation (TwoNN), Neighborhood overlap, PCA, Sparse Autoencoders (SAE)

GPT

The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024)measured

Intrinsic dimension expands then compresses; semantics peak at the minimum

Details

Valeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]

models: GPT-2-XL · method: Intrinsic dimension estimation (TwoNN), Neighborhood overlap, PCA, Sparse Autoencoders (SAE)
Shared Global and Local Geometry of Language Model Embeddings (2025)measured

Token-embedding orientation is shared within families but drops across them

Details

Lee et al. measure token-embedding geometry across the GPT-2, Llama-3, and Gemma-2 model families [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] Relative orientation/cosine structure is near-identical within a model family but drops sharply across families trained on different data [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] A k-NN-neighborhood PCA intrinsic-dimension estimator finds low-ID tokens form semantically coherent clusters while high-ID tokens do not [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] EMB2EMB, a linear least-squares (OLS) map fit over 100k shared tokens, transfers CAA-style steering vectors (refusal, sycophancy, corrigibility) between differently-sized models [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings]

models: GPT-2-Medium · method: PCA, Intrinsic dimension estimation (TwoNN), Cross-model direction transfer via ridge regression
Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021)measured

Contextualized embeddings lie on a low-dimensional manifold rising with depth

Details

Cai et al. estimate Local Intrinsic Dimension (k-NN expansion model, K=100) for BERT, DistilBERT, GPT, GPT-2, and ELMo on Penn Treebank [cai-etal-2021] Mean LID is far below ambient dimension in every model (BERT 5.6, DistilBERT 7.3, GPT 6.8, GPT-2 7.0, ELMo 9.1) versus 18.0-26.1 for static GloVe/word2vec embeddings measured the same way [cai-etal-2021] LID increases nearly linearly with layer depth across all models, a monotonic-growth trajectory unlike the expansion-then-compression profile TwoNN finds in protein/image models (a discrepancy left unreconciled) [cai-etal-2021] A qualitative PCA visualization shows GPT and GPT-2 later-layer embeddings resemble a "Swiss Roll" manifold thickening with depth, while BERT and DistilBERT do not [cai-etal-2021]

models: GPT-1 (OpenAI GPT), GPT-2-small · method: Intrinsic dimension estimation (TwoNN)

GPT-Neo

Memorization in Language Models through the Lens of Intrinsic Dimension (2025)measured

Lower input-sequence intrinsic dimension predicts more verbatim memorization

Details

Arnold treats each training text as a point cloud of BERT contextual embeddings and estimates its TwoNN intrinsic dimension, a property of the input sequence rather than a layer-wise hidden-state profile of the studied model [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Across GPT-Neo 125M/1.3B/2.7B and GPT-J-6B, lower-ID sequences are more likely to be verbatim-memorized and higher-ID sequences less likely, so intrinsic dimension acts as a suppressive signal for memorization [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] In the low-duplication regime, memorization declines inversely with intrinsic dimensionality across all model sizes [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] The relationship is purely observational: no detection classifier and no mutual-information analysis are performed [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Scope is restricted to exact-duplicate verbatim memorization under greedy decoding on 1,000 Pile sequences of 150 tokens [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension]

models: GPT-Neo-125M, GPT-Neo-1.3B, GPT-Neo-2.7B · method: Intrinsic dimension estimation (TwoNN)

GPT-J

Memorization in Language Models through the Lens of Intrinsic Dimension (2025)measured

Lower input-sequence intrinsic dimension predicts more verbatim memorization

Details

Arnold treats each training text as a point cloud of BERT contextual embeddings and estimates its TwoNN intrinsic dimension, a property of the input sequence rather than a layer-wise hidden-state profile of the studied model [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Across GPT-Neo 125M/1.3B/2.7B and GPT-J-6B, lower-ID sequences are more likely to be verbatim-memorized and higher-ID sequences less likely, so intrinsic dimension acts as a suppressive signal for memorization [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] In the low-duplication regime, memorization declines inversely with intrinsic dimensionality across all model sizes [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] The relationship is purely observational: no detection classifier and no mutual-information analysis are performed [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Scope is restricted to exact-duplicate verbatim memorization under greedy decoding on 1,000 Pile sequences of 150 tokens [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension]

models: GPT-J-6B · method: Intrinsic dimension estimation (TwoNN)

Mistral

A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension (2024)measured

In-context learning induces higher intrinsic dimension than fine-tuning

Details

Janapati & Ji apply TwoNN to last-token representations of Llama-3-8B, Llama-2-13B, Llama-2-7B, and Mistral-7B-v0.3 across 8 tasks, comparing ICL, LoRA fine-tuning, and zero-shot [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] ICL with k>=5 demonstrations induces a consistently higher intrinsic-dimension profile across all layers than SFT or zero-shot [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] This holds even though SFT reaches higher task accuracy (e.g. MMLU 0.542 vs ICL 0.531), dissociating performance from manifold dimensionality [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension] ID versus number of demonstrations is non-monotonic, rising then plateauing past k approximately 5-10 [janapati-ji-2024-learning-paradigms-via-intrinsic-dimension]

models: Mistral-7B-v0.3 · method: Intrinsic dimension estimation (TwoNN)
Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024)measured

Nonlinear intrinsic dimension phase-transitions with emergent zero-shot competence

Details

Lee et al. track per-layer nonlinear intrinsic dimension (TwoNN) and linear PCA effective dimension across pretraining of Pythia-410M/1.4B/6.9B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] On synthetic data both measures scale with dataset compositionality in fully-trained models, but nonlinear ID shows a sharp phase transition around step ~10^3 coinciding with the onset of zero-shot competence [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] PCA effective dimension tracks superficial/Kolmogorov complexity (gzip compressibility, early-training Spearman rho near 1.0) while TwoNN ID never correlates with gzip [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] The linear-superficial versus nonlinear-semantic dissociation generalizes to fully-trained Llama-3-8B and Mistral-7B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime]

models: Mistral-7B · method: Intrinsic dimension estimation (TwoNN), PCA
Rethinking Intrinsic Dimension Estimation in Neural Representations (2026)measured

The hump-shaped ID profile is likely a TwoNN estimator artifact

Details

Schulte & Rugamer prove that standard network layers (linear/conv, ReLU, softmax, pooling, residual, normalization, self-attention) are Lipschitz maps, so true pointwise and Hausdorff intrinsic dimension can only stay equal or decrease across a layer, never increase [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Yet the standard TwoNN/MLE/GRIDE pipeline on ResNet-34, Llama-3.1-8B, Mistral-7B-v0.3, and Pythia-6.9B reproduces the familiar hump-shaped expansion-then-compression ID profile, a pattern their theorem proves cannot be the true ID [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] Investigating what the estimators actually respond to (neighbor distances, ambient dimension, cosine similarity, norm, entropy), they conclude the mid-layer-ID-peak abstraction narrative is very likely a systematic estimator artifact [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation] The caution applies specifically to TwoNN and related nearest-neighbor-ratio estimators, not to MST-based, spectral-entropy/effective-rank, or SVD-based measures [schulte-rugamer-2026-rethinking-intrinsic-dimension-estimation]

models: Mistral-7B-v0.3 · method: Intrinsic dimension estimation (TwoNN)
Revisiting Hallucination Detection with Effective Rank-based Uncertainty (2025)measured

Multi-response effective rank detects hallucination at AUROC 0.84-0.86

Details

Wang et al. compute the effective rank (Roy-Vetterli spectral-entropy dimensionality) of a matrix of embeddings sampled across multiple generated responses and layers of Llama-2-7b-chat, Llama-2-13b-chat, and Mistral-7B-v0.1 [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Used as a hallucination signal it reaches AUROC around 0.84-0.86 across QA benchmarks, competitive with or exceeding semantic-entropy and self-consistency baselines [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] The 13B BioASQ configuration scores 0.8234, slightly below the headline 0.84-0.86 range [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Ablations vary the number of generations and the layer-selection strategy (middle layer vs last-5 layers) [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty]

models: Mistral-7B-v0.1 · method: Intrinsic dimension estimation (TwoNN)
Lines of Thought in Large Language Models (2024)measured

Token trajectories trace a curved low-dimensional Langevin manifold

Details

Sarfati et al. track each token's hidden-state trajectory through GPT-2-Medium's 24 layers (replicated on Llama-2-7B, Mistral-7B-v0.1, Llama-3.2-1B/3B) and find ensembles cluster on a low-dimensional curved manifold [sarfati-etal-2024-lines-of-thought] Truncating to the top ~256 of 1024 per-layer SVD dimensions preserves ~90% of the next-token distribution's information [sarfati-etal-2024-lines-of-thought] Layer-to-layer evolution follows a rotation-and-stretch law plus exponentially-growing Gaussian noise, a linear Langevin/Ornstein-Uhlenbeck SDE extracted from ensemble statistics [sarfati-etal-2024-lines-of-thought] Simulated SDE trajectories reproduce real statistics (a linear SVM separates them at near-chance 46-61%), while untrained-model trajectories move in straight parallel lines, so the structure is training-dependent [sarfati-etal-2024-lines-of-thought]

models: Mistral-7B-v0.1 · method: SVD, Geometric analysis

GPT-2

Representational Curvature Modulates Behavioral Uncertainty in Large Language Models (2026)measured

Trajectory-subspace curvature perturbations causally modulate next-token entropy

Details

King et al. define contextual curvature as a 3-token backward average of the angle between consecutive residual-stream displacement vectors, extending the Hosseini-Fedorenko straightening measure [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Across GPT-2 XL and Pythia-2.8B, contextual curvature predicts next-token entropy (Pearson r peaking ~0.15), rising across layers to a peak near the middle-layer curvature minimum [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] In Pythia's training the coupling is absent at 0-0.07% of a 300B-token run and emerges sharply around 0.7% [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Only perturbations restricted to the recent-displacement subspace or its 2D plane causally move entropy; random, random-subspace, activation-PCA, and full-space perturbations of matched magnitude do not [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty] Training a 12-layer model from scratch with a curvature-regularizing auxiliary loss lowers token-level entropy without degrading validation loss [king-etal-2026-representational-curvature-modulates-behavioral-uncertainty]

models: GPT-2 XL, Custom 12-layer GPT-2-Small (trained from scratch, 100M tokens) · method: Geometric analysis, Causal interventions (steering)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026)measured

A participation-ratio spectral signal finds attention circuits label-free across scale

Details

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

models: nanoGPT-style GPT-2 124M (FineWeb-10B) · method: Participation-ratio spectral signal (per-head effective rank over training), Causal interventions (steering)
Lines of Thought in Large Language Models (2024)measured

Token trajectories trace a curved low-dimensional Langevin manifold

Details

Sarfati et al. track each token's hidden-state trajectory through GPT-2-Medium's 24 layers (replicated on Llama-2-7B, Mistral-7B-v0.1, Llama-3.2-1B/3B) and find ensembles cluster on a low-dimensional curved manifold [sarfati-etal-2024-lines-of-thought] Truncating to the top ~256 of 1024 per-layer SVD dimensions preserves ~90% of the next-token distribution's information [sarfati-etal-2024-lines-of-thought] Layer-to-layer evolution follows a rotation-and-stretch law plus exponentially-growing Gaussian noise, a linear Langevin/Ornstein-Uhlenbeck SDE extracted from ensemble statistics [sarfati-etal-2024-lines-of-thought] Simulated SDE trajectories reproduce real statistics (a linear SVM separates them at near-chance 46-61%), while untrained-model trajectories move in straight parallel lines, so the structure is training-dependent [sarfati-etal-2024-lines-of-thought]

models: GPT-2 Medium · method: SVD, Geometric analysis

OPT

Abstraction Induces the Brain Alignment of Language and Speech Models (2026)measured

Layer intrinsic-dimension peak aligns with the best brain-predicting layer

Details

Cheng et al. estimate per-layer intrinsic dimension via the GRIDE estimator in OPT (125M/1.3B/13B), Pythia (160M/410M/6.9B), WavLM (base-plus, large), and Whisper (large encoder) [cheng-etal-2026-abstraction-brain-alignment] Layerwise intrinsic dimension correlates with how well each layer predicts real human brain responses (fMRI rho=0.76, ECoG rho=0.43, both p<0.05) [cheng-etal-2026-abstraction-brain-alignment] The ID-peak layer and the best-brain-predicting layer are usually within 0-1 layers of each other [cheng-etal-2026-abstraction-brain-alignment] Brain-tuning WavLM's best layer (layer 9) to predict fMRI causally raises both its intrinsic dimension and its semantic content, while a random-Fourier-features control shows raised ID alone is insufficient [cheng-etal-2026-abstraction-brain-alignment]

models: OPT-125M, OPT-1.3B, OPT-13B · method: Intrinsic dimension estimation (TwoNN), Causal interventions (steering)

WavLM

Abstraction Induces the Brain Alignment of Language and Speech Models (2026)measured

Layer intrinsic-dimension peak aligns with the best brain-predicting layer

Details

Cheng et al. estimate per-layer intrinsic dimension via the GRIDE estimator in OPT (125M/1.3B/13B), Pythia (160M/410M/6.9B), WavLM (base-plus, large), and Whisper (large encoder) [cheng-etal-2026-abstraction-brain-alignment] Layerwise intrinsic dimension correlates with how well each layer predicts real human brain responses (fMRI rho=0.76, ECoG rho=0.43, both p<0.05) [cheng-etal-2026-abstraction-brain-alignment] The ID-peak layer and the best-brain-predicting layer are usually within 0-1 layers of each other [cheng-etal-2026-abstraction-brain-alignment] Brain-tuning WavLM's best layer (layer 9) to predict fMRI causally raises both its intrinsic dimension and its semantic content, while a random-Fourier-features control shows raised ID alone is insufficient [cheng-etal-2026-abstraction-brain-alignment]

models: WavLM-base-plus, WavLM Large · method: Intrinsic dimension estimation (TwoNN), Causal interventions (steering)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models (2026)measured

LID separates adversarial from benign speech perturbations, transcript-free

Details

Arcos-Holzinger et al. compute per-layer Local Intrinsic Dimensionality (framework GRIDS) on WavLM-base and wav2vec 2.0 base representations under acoustic perturbation [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models] LID rises for all low-SNR perturbations, but as SNR increases benign noise's LID converges back toward the clean profile while adversarial inputs retain elevated early-layer LID [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models] This per-layer LID divergence distinguishes adversarial from benign perturbations at AUROC 0.78-1.00 without ground-truth transcripts [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models]

models: WavLM Base · method: Local Intrinsic Dimensionality (LID) under perturbation

Whisper

Abstraction Induces the Brain Alignment of Language and Speech Models (2026)measured

Layer intrinsic-dimension peak aligns with the best brain-predicting layer

Details

Cheng et al. estimate per-layer intrinsic dimension via the GRIDE estimator in OPT (125M/1.3B/13B), Pythia (160M/410M/6.9B), WavLM (base-plus, large), and Whisper (large encoder) [cheng-etal-2026-abstraction-brain-alignment] Layerwise intrinsic dimension correlates with how well each layer predicts real human brain responses (fMRI rho=0.76, ECoG rho=0.43, both p<0.05) [cheng-etal-2026-abstraction-brain-alignment] The ID-peak layer and the best-brain-predicting layer are usually within 0-1 layers of each other [cheng-etal-2026-abstraction-brain-alignment] Brain-tuning WavLM's best layer (layer 9) to predict fMRI causally raises both its intrinsic dimension and its semantic content, while a random-Fourier-features control shows raised ID alone is insufficient [cheng-etal-2026-abstraction-brain-alignment]

models: Whisper large · method: Intrinsic dimension estimation (TwoNN), Causal interventions (steering)

wav2vec 2.0

Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models (2026)measured

LID separates adversarial from benign speech perturbations, transcript-free

Details

Arcos-Holzinger et al. compute per-layer Local Intrinsic Dimensionality (framework GRIDS) on WavLM-base and wav2vec 2.0 base representations under acoustic perturbation [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models] LID rises for all low-SNR perturbations, but as SNR increases benign noise's LID converges back toward the clean profile while adversarial inputs retain elevated early-layer LID [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models] This per-layer LID divergence distinguishes adversarial from benign perturbations at AUROC 0.78-1.00 without ground-truth transcripts [arcosholzinger-etal-2026-dimensionality-aware-anomaly-detection-in-ssl-speech-models]

models: wav2vec2-base-960h · method: Local Intrinsic Dimensionality (LID) under perturbation

ConvNeXt

Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and the Human Brain (2026)measured

Lower local intrinsic dimension predicts vision-model alignment and generalization

Details

Yu et al. estimate local intrinsic dimension via the Levina-Bickel maximum-likelihood kNN estimator across a pool of 91 vision models (ConvNeXt, ResNet, ResMLP, ViT), with a 51-model architecture-balanced subset for cross-architecture comparison [yu-etal-2026-local-intrinsic-dimension-alignment] Local ID is significantly negatively correlated with AI-AI representational alignment, AI-brain alignment (fMRI, Natural Scenes Dataset), and ImageNet-1K generalization [yu-etal-2026-local-intrinsic-dimension-alignment] The correlation is strongest at small (local) neighborhood scale K and weakens toward global estimates, with a matched-subsample control confirming genuine local structure [yu-etal-2026-local-intrinsic-dimension-alignment] Increasing model capacity and training-data scale systematically reduces local ID (local ID vs log-parameters Corr=-0.756, p<0.001) [yu-etal-2026-local-intrinsic-dimension-alignment] PCA is used only as a top-300-component control to equalize ambient dimensionality, not as the ID estimator, and robustness is checked with the MOM and MADA estimators [yu-etal-2026-local-intrinsic-dimension-alignment]

models: ConvNeXt (image classifier, various sizes, ImageNet) · method: Local Intrinsic Dimensionality (LID) under perturbation, PCA

Vision Transformer (ViT)

Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and the Human Brain (2026)measured

Lower local intrinsic dimension predicts vision-model alignment and generalization

Details

Yu et al. estimate local intrinsic dimension via the Levina-Bickel maximum-likelihood kNN estimator across a pool of 91 vision models (ConvNeXt, ResNet, ResMLP, ViT), with a 51-model architecture-balanced subset for cross-architecture comparison [yu-etal-2026-local-intrinsic-dimension-alignment] Local ID is significantly negatively correlated with AI-AI representational alignment, AI-brain alignment (fMRI, Natural Scenes Dataset), and ImageNet-1K generalization [yu-etal-2026-local-intrinsic-dimension-alignment] The correlation is strongest at small (local) neighborhood scale K and weakens toward global estimates, with a matched-subsample control confirming genuine local structure [yu-etal-2026-local-intrinsic-dimension-alignment] Increasing model capacity and training-data scale systematically reduces local ID (local ID vs log-parameters Corr=-0.756, p<0.001) [yu-etal-2026-local-intrinsic-dimension-alignment] PCA is used only as a top-300-component control to equalize ambient dimensionality, not as the ID estimator, and robustness is checked with the MOM and MADA estimators [yu-etal-2026-local-intrinsic-dimension-alignment]

models: Vision Transformer (ViT, image classifier, various sizes, ImageNet) · method: Local Intrinsic Dimensionality (LID) under perturbation, PCA

ResMLP

Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and the Human Brain (2026)measured

Lower local intrinsic dimension predicts vision-model alignment and generalization

Details

Yu et al. estimate local intrinsic dimension via the Levina-Bickel maximum-likelihood kNN estimator across a pool of 91 vision models (ConvNeXt, ResNet, ResMLP, ViT), with a 51-model architecture-balanced subset for cross-architecture comparison [yu-etal-2026-local-intrinsic-dimension-alignment] Local ID is significantly negatively correlated with AI-AI representational alignment, AI-brain alignment (fMRI, Natural Scenes Dataset), and ImageNet-1K generalization [yu-etal-2026-local-intrinsic-dimension-alignment] The correlation is strongest at small (local) neighborhood scale K and weakens toward global estimates, with a matched-subsample control confirming genuine local structure [yu-etal-2026-local-intrinsic-dimension-alignment] Increasing model capacity and training-data scale systematically reduces local ID (local ID vs log-parameters Corr=-0.756, p<0.001) [yu-etal-2026-local-intrinsic-dimension-alignment] PCA is used only as a top-300-component control to equalize ambient dimensionality, not as the ID estimator, and robustness is checked with the MOM and MADA estimators [yu-etal-2026-local-intrinsic-dimension-alignment]

models: ResMLP (image classifier, various sizes, ImageNet) · method: Local Intrinsic Dimensionality (LID) under perturbation, PCA

BERT

Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021)measured

Contextualized embeddings lie on a low-dimensional manifold rising with depth

Details

Cai et al. estimate Local Intrinsic Dimension (k-NN expansion model, K=100) for BERT, DistilBERT, GPT, GPT-2, and ELMo on Penn Treebank [cai-etal-2021] Mean LID is far below ambient dimension in every model (BERT 5.6, DistilBERT 7.3, GPT 6.8, GPT-2 7.0, ELMo 9.1) versus 18.0-26.1 for static GloVe/word2vec embeddings measured the same way [cai-etal-2021] LID increases nearly linearly with layer depth across all models, a monotonic-growth trajectory unlike the expansion-then-compression profile TwoNN finds in protein/image models (a discrepancy left unreconciled) [cai-etal-2021] A qualitative PCA visualization shows GPT and GPT-2 later-layer embeddings resemble a "Swiss Roll" manifold thickening with depth, while BERT and DistilBERT do not [cai-etal-2021]

models: BERT-base-uncased · method: Intrinsic dimension estimation (TwoNN)

DistilBERT

Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021)measured

Contextualized embeddings lie on a low-dimensional manifold rising with depth

Details

Cai et al. estimate Local Intrinsic Dimension (k-NN expansion model, K=100) for BERT, DistilBERT, GPT, GPT-2, and ELMo on Penn Treebank [cai-etal-2021] Mean LID is far below ambient dimension in every model (BERT 5.6, DistilBERT 7.3, GPT 6.8, GPT-2 7.0, ELMo 9.1) versus 18.0-26.1 for static GloVe/word2vec embeddings measured the same way [cai-etal-2021] LID increases nearly linearly with layer depth across all models, a monotonic-growth trajectory unlike the expansion-then-compression profile TwoNN finds in protein/image models (a discrepancy left unreconciled) [cai-etal-2021] A qualitative PCA visualization shows GPT and GPT-2 later-layer embeddings resemble a "Swiss Roll" manifold thickening with depth, while BERT and DistilBERT do not [cai-etal-2021]

models: DistilBERT-base-uncased · method: Intrinsic dimension estimation (TwoNN)

ELMo

Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021)measured

Contextualized embeddings lie on a low-dimensional manifold rising with depth

Details

Cai et al. estimate Local Intrinsic Dimension (k-NN expansion model, K=100) for BERT, DistilBERT, GPT, GPT-2, and ELMo on Penn Treebank [cai-etal-2021] Mean LID is far below ambient dimension in every model (BERT 5.6, DistilBERT 7.3, GPT 6.8, GPT-2 7.0, ELMo 9.1) versus 18.0-26.1 for static GloVe/word2vec embeddings measured the same way [cai-etal-2021] LID increases nearly linearly with layer depth across all models, a monotonic-growth trajectory unlike the expansion-then-compression profile TwoNN finds in protein/image models (a discrepancy left unreconciled) [cai-etal-2021] A qualitative PCA visualization shows GPT and GPT-2 later-layer embeddings resemble a "Swiss Roll" manifold thickening with depth, while BERT and DistilBERT do not [cai-etal-2021]

models: ELMo (AllenNLP biLM, 1B Word Benchmark) · method: Intrinsic dimension estimation (TwoNN)

DINOv2

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension (2026)measured

MST intrinsic dimension of frozen SSL features predicts probe accuracy

Details

Mordacq et al. estimate intrinsic dimension of frozen penultimate-layer features from 33 pretrained self-supervised vision checkpoints across 14 methods (VICReg, DINO, BarlowTwins, DINOv2, DINOv3, I-JEPA, iBOT, CLIP, EVA-CLIP, SigLIP, PE-Core, Franca) using a minimum-spanning-tree estimator (IdEst) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] MST intrinsic dimension is strongly negatively correlated with downstream linear-probe accuracy across four datasets (Spearman rho approximately -0.8, e.g. ImageNet -0.74, iNat-18 -0.84; Kendall tau approximately -0.6) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] Lower-dimensional frozen representations transfer better, making intrinsic dimension a label-free proxy for representation quality without a linear-probe run [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] The relationship is correlational, over frozen features and linear-probe evaluation only [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension]

models: DINOv2 ViT-B/14 · method: MST-based intrinsic dimension estimator

iBOT

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension (2026)measured

MST intrinsic dimension of frozen SSL features predicts probe accuracy

Details

Mordacq et al. estimate intrinsic dimension of frozen penultimate-layer features from 33 pretrained self-supervised vision checkpoints across 14 methods (VICReg, DINO, BarlowTwins, DINOv2, DINOv3, I-JEPA, iBOT, CLIP, EVA-CLIP, SigLIP, PE-Core, Franca) using a minimum-spanning-tree estimator (IdEst) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] MST intrinsic dimension is strongly negatively correlated with downstream linear-probe accuracy across four datasets (Spearman rho approximately -0.8, e.g. ImageNet -0.74, iNat-18 -0.84; Kendall tau approximately -0.6) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] Lower-dimensional frozen representations transfer better, making intrinsic dimension a label-free proxy for representation quality without a linear-probe run [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] The relationship is correlational, over frozen features and linear-probe evaluation only [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension]

models: iBOT ViT-B/16 · method: MST-based intrinsic dimension estimator

CLIP (Contrastive Language-Image Pretraining)

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension (2026)measured

MST intrinsic dimension of frozen SSL features predicts probe accuracy

Details

Mordacq et al. estimate intrinsic dimension of frozen penultimate-layer features from 33 pretrained self-supervised vision checkpoints across 14 methods (VICReg, DINO, BarlowTwins, DINOv2, DINOv3, I-JEPA, iBOT, CLIP, EVA-CLIP, SigLIP, PE-Core, Franca) using a minimum-spanning-tree estimator (IdEst) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] MST intrinsic dimension is strongly negatively correlated with downstream linear-probe accuracy across four datasets (Spearman rho approximately -0.8, e.g. ImageNet -0.74, iNat-18 -0.84; Kendall tau approximately -0.6) [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] Lower-dimensional frozen representations transfer better, making intrinsic dimension a label-free proxy for representation quality without a linear-probe run [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension] The relationship is correlational, over frozen features and linear-probe evaluation only [mordacq-etal-2026-idest-assessing-ssl-representations-via-intrinsic-dimension]

models: CLIP ViT-B/32 · method: MST-based intrinsic dimension estimator

Custom Research Transformer (small, purpose-built for interpretability studies)

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026)measured

A participation-ratio spectral signal finds attention circuits label-free across scale

Details

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

models: TS-51M (custom 8-layer x 512d x 16-head transformer, TinyStories, 6 pretraining seeds) · method: Participation-ratio spectral signal (per-head effective rank over training), Causal interventions (steering)

OLMo

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026)measured

A participation-ratio spectral signal finds attention circuits label-free across scale

Details

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

models: OLMo-1B · method: Participation-ratio spectral signal (per-head effective rank over training), Causal interventions (steering)

OLMoE

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026)measured

A participation-ratio spectral signal finds attention circuits label-free across scale

Details

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

models: OLMoE-1B-7B · method: Participation-ratio spectral signal (per-head effective rank over training), Causal interventions (steering)

BLOOM

The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024)measured

Embedding intrinsic dimension expands early in pretraining, then compresses

Details

Razzhigaev et al. apply TwoNN (cross-validated against Manifold-Adaptive Dimension Estimation and the Method of Moments) to embeddings sampled across real pretraining checkpoints of Bloom-3B and Pythia-2.8B [razzhigaev-etal-2024-shape-of-learning] Intrinsic dimension rises in the initial phase of training, then compresses toward the end, a training-time (not depth-wise) dimensionality trajectory [razzhigaev-etal-2024-shape-of-learning] The analysis is purely observational, with no causal intervention [razzhigaev-etal-2024-shape-of-learning]

models: BLOOM-3B · method: Intrinsic dimension estimation (TwoNN)

Qwen

The Geometry of Reasoning: Flowing Logics in Representation Space (2026)measured

Reasoning-step curvature tracks logical structure, not surface form

Details

Zhou et al. model each reasoning step as a mean-pooled final-layer hidden state and compute the Menger curvature of the step sequence across 2,430 sequences that vary logic, topic, and language independently [zhou-etal-2026-geometry-of-reasoning] Curvature similarity is consistently higher for logic-matched pairs than topic- or language-matched pairs (e.g. Qwen3-0.6B: logic 0.53 vs topic 0.11 vs language 0.13; Llama3-8B: 0.58 vs 0.13 vs 0.17) [zhou-etal-2026-geometry-of-reasoning] Raw position similarity is instead dominated by language (Qwen3-0.6B: language 0.85 vs logic 0.26), so second-order curvature, not position, is organized by logical structure [zhou-etal-2026-geometry-of-reasoning] A step-shuffle control collapses velocity and curvature similarity toward zero while leaving position similarity high, and no causal intervention is performed [zhou-etal-2026-geometry-of-reasoning]

models: Qwen1.5-0.5B, Qwen2-0.5B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B · method: Geometric analysis, PCA
Reasoning emerges from constrained inference manifolds in large language models (2026)measured

Chain-of-thought trajectories collapse to under ten intrinsic dimensions

Details

Ma et al. estimate intrinsic dimension with the TLE estimator (k=20) for the static vocabulary-embedding cloud (D_world) and for autoregressive last-token chain-of-thought trajectories (D_stim) across Qwen2.5, Qwen3, Gemma 3, and DeepSeek-R1-Distill-Qwen families [ma-etal-2026-constrained-inference-manifolds] Deep-layer reasoning trajectories collapse onto a manifold of fewer than ten intrinsic degrees of freedom despite thousands of ambient dimensions, while the vocabulary cloud stays near ambient dimension [ma-etal-2026-constrained-inference-manifolds] A composite reasoning-health diagnostic H = log(D_world)*V/exp(0.1*D_stim), combining an information-volume term V, correlates with benchmark scores at Spearman rho>0.9 [ma-etal-2026-constrained-inference-manifolds] Token-shuffle, non-cognitive-prompt, and truncation controls rule out trivial confounds, and no causal intervention is performed [ma-etal-2026-constrained-inference-manifolds]

models: Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-14B, Qwen2.5-32B, Qwen2.5-72B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3-32B · method: Intrinsic dimension estimation (TwoNN), Geometric analysis
Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs (2026)measured

Evaluation-awareness probe depth shifts late-to-early with model scale

Details

Manek projects a diff-of-means evaluation-awareness direction (from 203 contrastive prompt pairs) onto residual-stream activations at every layer across 11 open-weight models (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] The relative depth at which per-layer decoding AUROC peaks shifts from late layers in small models to the earliest layers in large ones (Qwen2.5 1.5B/3B peak at 0.96-0.97 vs 14B/32B at 0.021-0.031) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Gemma2 shows the same late-to-early shift (0.885 to 0.304) while Llama-3.2 stays mid-layer across both tested sizes [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Peak AUROC ranges 0.586-0.873 and is non-monotonic with scale within a family [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale]

models: Qwen2.5-0.5B-Instruct, Qwen2.5-1.5B, Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-32B · method: Direction Extraction, Linear probing, Geometric analysis

DeepSeek

Reasoning emerges from constrained inference manifolds in large language models (2026)measured

Chain-of-thought trajectories collapse to under ten intrinsic dimensions

Details

Ma et al. estimate intrinsic dimension with the TLE estimator (k=20) for the static vocabulary-embedding cloud (D_world) and for autoregressive last-token chain-of-thought trajectories (D_stim) across Qwen2.5, Qwen3, Gemma 3, and DeepSeek-R1-Distill-Qwen families [ma-etal-2026-constrained-inference-manifolds] Deep-layer reasoning trajectories collapse onto a manifold of fewer than ten intrinsic degrees of freedom despite thousands of ambient dimensions, while the vocabulary cloud stays near ambient dimension [ma-etal-2026-constrained-inference-manifolds] A composite reasoning-health diagnostic H = log(D_world)*V/exp(0.1*D_stim), combining an information-volume term V, correlates with benchmark scores at Spearman rho>0.9 [ma-etal-2026-constrained-inference-manifolds] Token-shuffle, non-cognitive-prompt, and truncation controls rule out trivial confounds, and no causal intervention is performed [ma-etal-2026-constrained-inference-manifolds]

models: DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-14B, DeepSeek-R1-Distill-Qwen-32B · method: Intrinsic dimension estimation (TwoNN), Geometric analysis

RoBERTa

Less is More: Local Intrinsic Dimensions of Contextual Language Models (2025)measured

Local dimension of real fine-tuned RoBERTa embeddings drops with successful dialogue-state tracking and rises with overfitting onset on real emotion-recognition fine-tuning

Details

The TwoNN local-intrinsic-dimension estimator is applied to the 768-dimensional last-layer embeddings of a real RoBERTa-base encoder, both as a masked-LM base model and after real task fine-tuning: a TripPy-R dialogue-state tracker fine-tuned on the real MultiWOZ 2.1 dataset, and a separate fine-tune on the real EmoWOZ 7-class emotion-recognition dataset [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models] Local dimension estimates for the MultiWOZ-fine-tuned RoBERTa are markedly lower than for the base (un-fine-tuned) model, and this dimension drop coincides with the onset of rising validation accuracy in an auxiliary synthetic modular-arithmetic grokking task used to validate the estimator against a known transition point [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models] On the real EmoWOZ fine-tune, local dimension instead rises after training epoch 1, at the same point validation loss begins to increase, tying rising local dimension to the onset of overfitting [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models]

models: RoBERTa-base (TripPy-R dialogue-state tracker, fine-tuned on MultiWOZ 2.1), RoBERTa-base (fine-tuned on EmoWOZ, 7-class emotion recognition) · method: Local Intrinsic Dimensionality (LID) under perturbation

scGPT

Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals (2026)measured

scGPT attention weights encode a compact ~8-10D hematopoietic manifold

Details

Kendiukhov exports a fixed operator from scGPT's frozen attention value-projection weights (no forward-pass activations) and compresses it into a latent space whose effective dimensionality plateaus at ~8-10 [kendiukhov-2026-scgpt-hematopoietic-manifold] The manifold's geodesic distances correlate with an independent hematopoietic developmental ontology (erythroid rho=0.768 p=0.0017, granulocytic rho=0.568, trunk rho=0.611; internal rho=0.835) [kendiukhov-2026-scgpt-hematopoietic-manifold] It validates on a strict non-overlap external panel (Tabula Sapiens, 564,253 cells) and zero-shot transfer to an immune panel (trustworthiness 0.993, blocked-permutation p=0.0005) [kendiukhov-2026-scgpt-hematopoietic-manifold] A lightweight readout on the extracted manifold beats scVI, Palantir, DPT, CellTypist, and PCA on pseudotime ordering (|rho|=0.439 vs 0.331) with ~1,000x fewer parameters [kendiukhov-2026-scgpt-hematopoietic-manifold]

models: scGPT (whole-human pretrained checkpoint) · method: Geometric analysis
Multi-Dimensional Spectral Geometry of Biological Knowledge in Single-Cell Transformer Representations (2026)measured

scGPT gene-embedding effective rank collapses 14-fold across layers

Details

Kendiukhov performs per-layer SVD on scGPT's gene-embedding matrix across its 12 layers; effective rank collapses monotonically 23.6 to 1.6 (Spearman rho=-1.000) and the top singular vector's variance fraction rises 53.7% to 93.4% [kendiukhov-2026-scgpt-spectral-geometry] TwoNN intrinsic dimension falls 32.6 to 18.1 and the participation ratio drops 6.1-fold, while a feature-shuffle control rebounds effective rank to 28.9, confirming a genuine (non-artifactual) collapse [kendiukhov-2026-scgpt-spectral-geometry] The final compressed layer's low-rank subspace matches independent biology: STRING PPI co-pole rate 0.226 vs 0.124 null, TRRUST TF-target AUROC up to 0.789, and cell-type marker AUROC 0.851 vs 0.488 chance [kendiukhov-2026-scgpt-spectral-geometry] Extensive confound controls reject persistent-homology and feed-forward-loop explanations, and no causal intervention is performed [kendiukhov-2026-scgpt-spectral-geometry]

models: scGPT (whole-human pretrained checkpoint) · method: Intrinsic dimension estimation (TwoNN)