MATH · IN · MODELS

A fitted linear translation map recovers apparently-deleted SAE concept latents across continual-learning checkpoints

measured in 1 paper

Filus, Faber, Corizzo & Kanan train a BatchTopK sparse autoencoder per task on frozen real ResNet-18 activations (via the Mammoth continual-learning framework) on 2seq-CIFAR10, 2seq-tiny-ImageNet, and 10seq-tiny-ImageNet, then check which task-t SAE latents stop firing after continual learning on later tasks ("apparent deletion") [filus-etal-2026-lost-or-hidden-concept-level-forgetting-supervised-continual-learning] A least-squares linear map T (with bias term) translating post-continual-learning frozen features back toward the earlier task's representation space is fit per checkpoint pair, and re-running the frozen earlier-task SAE on translated features recovers many apparently-deleted latents, distinguishing concepts that are genuinely lost from concepts that are merely hidden behind a recoverable linear reparameterization [filus-etal-2026-lost-or-hidden-concept-level-forgetting-supervised-continual-learning] A nonlinear MLP translator gives only marginal additional recovery over the linear map, and deletion ratios are highest under naive SGD and EWC continual-learning strategies and lowest under DER++ and LwF, with concept decodability (via a separate logistic-regression probe) degrading further as more tasks accumulate [filus-etal-2026-lost-or-hidden-concept-level-forgetting-supervised-continual-learning]

Method

Papers

Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning — Filus, Katarzyna, Faber, Kamil, Corizzo, Roberto, Kanan, Christopher2026 · arXiv:2605.16374