An encoding probe decomposes feature-group variance in wav2vec2 and BERT
measured in 1 paperShen et al. fit ridge-regression "encoding probes" that reconstruct wav2vec2-base and BERT-base activations from interpretable feature sets (acoustics, phonetics, speaker-identity, syntax, lexicon), reporting unexplained variance (1 minus R-squared) under feature-group ablation [shen-etal-2026-beyond-decodability-reconstructing-language-model-representations-with-an-encoding-probe] Speaker-identity's contribution to explained variance shifts substantially depending on whether wav2vec2 is left self-supervised or fine-tuned for ASR versus speaker-ID [shen-etal-2026-beyond-decodability-reconstructing-language-model-representations-with-an-encoding-probe] Syntactic and lexical feature groups contribute largely independently and additively, with unexplained-variance differences stable to within 0.002 across seeds [shen-etal-2026-beyond-decodability-reconstructing-language-model-representations-with-an-encoding-probe] No causal intervention is performed; the analysis is explicitly observational [shen-etal-2026-beyond-decodability-reconstructing-language-model-representations-with-an-encoding-probe]