methods / Causal Validation / Training-loss uniformity/alignment regularization
Training-loss uniformity/alignment regularization
Adds explicit uniformity and/or alignment loss terms to a contrastive objective and re-trains or fine-tunes the model, then measures the resulting change in global geometric statistics (dimensionality, centroid separation) and downstream task performance - a training-time causal manipulation of the whole representation space, rather than a post-hoc activation edit or a single steering vector.