AlphaZero's novel chess concepts are causally steerable and teachable
measured in 1 paperSchut et al. extract new (not human-labeled) chess concepts from AlphaZero's activations as sparse linear directions via L1-regularized convex optimization [schut-etal-2023-bridging-human-ai-knowledge-gap] Concepts are filtered for teachability (a lesser-trained student AZ can reproduce them) and novelty (SVD reconstruction error against the human-game activation span) [schut-etal-2023-bridging-human-ai-knowledge-gap] Adding a concept direction back via norm-matched interpolation causally improves AZ's puzzle-solving specifically on that concept's puzzles [schut-etal-2023-bridging-human-ai-knowledge-gap] Four human grandmasters shown AZ's concept-instantiating lines improve their own puzzle-solving, evidencing the concepts are learnable outside the network [schut-etal-2023-bridging-human-ai-knowledge-gap]