FFN value-vector directions causally suppress toxicity
measured in 1 paperGeva et al. decompose each FFN layer's additive update to the output distribution into per-parameter sub-updates, each a value-vector direction projected through the unembedding [geva-etal-2022-ffn-promoting-concepts] In GPT-2, suppressing value-vector directions that promote toxic-concept tokens causally reduces generated toxicity by roughly 50% [geva-etal-2022-ffn-promoting-concepts] An early-exit rule built on the same decomposition saves about 20% of inference compute on average [geva-etal-2022-ffn-promoting-concepts] This is a geometry-tied causal intervention on a specific vocabulary-space direction, refining the FFN key-value-memory framing [geva-etal-2022-ffn-promoting-concepts]
Structure
Context
FFN value-vector decomposition, vocabulary-space projection, causal toxicity suppression
Confirmed in models
Method
Papers
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space — Geva, Mor, Caciularu, Avi, Wang, Kevin Ro, Goldberg, Yoav