MATH · IN · MODELS

Local dimension of real fine-tuned RoBERTa embeddings drops with successful dialogue-state tracking and rises with overfitting onset on real emotion-recognition fine-tuning

measured in 1 paper

The TwoNN local-intrinsic-dimension estimator is applied to the 768-dimensional last-layer embeddings of a real RoBERTa-base encoder, both as a masked-LM base model and after real task fine-tuning: a TripPy-R dialogue-state tracker fine-tuned on the real MultiWOZ 2.1 dataset, and a separate fine-tune on the real EmoWOZ 7-class emotion-recognition dataset [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models] Local dimension estimates for the MultiWOZ-fine-tuned RoBERTa are markedly lower than for the base (un-fine-tuned) model, and this dimension drop coincides with the onset of rising validation accuracy in an auxiliary synthetic modular-arithmetic grokking task used to validate the estimator against a known transition point [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models] On the real EmoWOZ fine-tune, local dimension instead rises after training epoch 1, at the same point validation loss begins to increase, tying rising local dimension to the onset of overfitting [ruppik-etal-2025-local-intrinsic-dimensions-contextual-language-models]

Context

local intrinsic dimension, fine-tuning dynamics, overfitting detection, dialogue-state tracking

Papers

Less is More: Local Intrinsic Dimensions of Contextual Language Models — Ruppik, Benjamin Matthias, von Rohrscheidt, Julius, van Niekerk, Carel, Heck, Michael, Vukovic, Renato, Feng, Shutong, Lin, Hsien-chin, Lubis, Nurul, Rieck, Bastian, Zibrowius, Marcus, Gasic, Milica2025 · arXiv:2506.01034