MATH · IN · MODELS

A per-token mean-difference direction fingerprints narrow fine-tuning

measured in 1 paper

Minder et al. extract a diff-in-means direction between base and narrowly-fine-tuned activations across 33 fine-tuning organisms spanning 7 models from 1B to 32B [minder-etal-2026-narrow-finetuning-traces-in-activation-differences] Steering by adding this direction during generation produces text with high embedding-similarity to the actual fine-tuning corpus [minder-etal-2026-narrow-finetuning-traces-in-activation-differences] An LLM interpretability agent (GPT-5) given the direction correctly identifies the fine-tuning objective in 91% of organisms (30/33) [minder-etal-2026-narrow-finetuning-traces-in-activation-differences] That is more than twice as well at broad-objective identification and over 30x better at fine-grained detail than the best black-box baseline [minder-etal-2026-narrow-finetuning-traces-in-activation-differences]

Context

narrow-finetuning fingerprint, per-token diff-in-means, agentic fine-tuning-objective identification

Papers

Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — Minder, Julian, Dumas, Clément, Slocum, Stewart, Casademunt, Helena, Holmes, Cameron, West, Robert, Nanda, Neel2026 · arXiv:2510.13900