MATH · IN · MODELS
methods / Direction Extraction / Mahalanobis-cosine probe alignment

Mahalanobis-cosine probe alignment

Techniqueadvanced

Compares two probe-derived directions using cosine similarity computed in a Mahalanobis-whitened (covariance-normalized) coordinate frame rather than the raw activation space, giving a far stronger predictor of cross-domain probe generalization than standard cosine similarity.

Used in (1 observation)

structure: Linear Direction · models: Llama 3.3 70B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct · paper: The Truthfulness Spectrum Hypothesis