MATH · IN · MODELS
methods / Causal Validation / Subspace noise perturbation

Subspace noise perturbation

Techniqueintermediate

Injects Gaussian noise into a located low-dimensional subspace (rather than ablating or adding a fixed direction) and measures downstream accuracy degradation, comparing against noise of the same magnitude in a random subspace or the full activation space.

Used in (1 observation)

structure: Circle, Cone, Platonic Representation Hypothesis · models: GPT-2-small, Mistral-7B, Llama-3-8B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-8B · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, Not All Language Model Features Are One-Dimensionally Linear, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Do Sparse Autoencoders Capture Concept Manifolds?