MATH · IN · MODELS
methods / Causal Validation / Probe-Gradient Perturbation

Probe-Gradient Perturbation

Techniqueintermediate

Iteratively perturbs a single instance's hidden state using the gradient of an already-trained probe's loss with respect to that hidden state, pushing the representation toward (gradient descent) or away from (gradient ascent) the probed property, then measures the effect on downstream model behavior.

Used in (1 observation)

structure: Affine Subspace · models: DeBERTa-v2-xxlarge, GPT-Neo-1.3B · paper: More than Correlation: Do Large Language Models Learn Causal Representations of Space?