MATH · IN · MODELS

ICL and fine-tuning share a layer-17 transition but differ before it

measured in 1 paper

Doimo et al. apply density-peak clustering and intrinsic-dimension estimation to Llama3-8B last-token representations, comparing in-context learning and supervised fine-tuning layer by layer [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] Both regimes undergo a sharp two-phase transition around layer 17, marked by an ID peak and a jump in density-peak cluster geometry [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] Before the transition, ICL organizes representations into far more clusters (60-70 versus under 40) and more sharply separated ones (core-point fraction ~0.6) than SFT [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning] After the transition, SFT develops sharper probability modes encoding answer identity, showing the two regimes induce measurably different cluster geometry within the same base model [doimo-etal-2024-the-representation-landscape-of-few-shot-learning-and-fine-tuning]

Context

in-context learning, fine-tuning, density-peak clustering, layer-wise phase transition, training regime comparison

Confirmed in models

Papers

The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language Models — Doimo, Diego, Serra, Alessandro, Ansuini, Alessio, Cazzaniga, Alberto2024 · arXiv:2409.03662