Fine-tuning widens convex-cluster margins without rearranging the geometry
measured in 1 paperZhou & Srikumar reuse DirectProbe on five BERT checkpoints across four linearly-separable tasks to test what fine-tuning does to convex-cluster geometry [zhou-srikumar-2021-finetuning] Fine-tuning consistently increases the minimum inter-cluster margin between every label pair, the geometric mechanism behind accuracy gains [zhou-srikumar-2021-finetuning] A new Spatial Similarity metric shows this is not arbitrary rearrangement: pre/post cluster geometry retains Pearson correlation above 0.5 in higher layers [zhou-srikumar-2021-finetuning] The same metric explains the one exception (BERT-small preposition-supersense accuracy dropping) via the lowest train/test spatial similarity of any condition (0.44) [zhou-srikumar-2021-finetuning]
Structure
Context
convex hull, fine-tuning dynamics, representational stability, version space
Confirmed in models
Papers
A Closer Look at How Fine-tuning Changes BERT — Zhou, Yichu, Srikumar, Vivek