methods / Causal Validation / Causal Tracing
Causal Tracing
Corrupts a subject's token embeddings with Gaussian noise, then restores one clean hidden state at a time during the corrupted run, measuring how much each individual state's restoration recovers the original prediction — a noising-and-selective-restoration variant of activation patching, used to localize which layer and token position causally mediates a specific factual prediction.