MATH · IN · MODELS

Boundless DAS scales causal-alignment search to a 7B LLM

measured in 1 paper

Wu et al. scale Distributed Alignment Search to Alpaca-7B on a "price tagging" task using Boundless DAS: a sigmoid-parameterized differentiable boundary that learns the target subspace's dimensionality by gradient descent instead of a brute-force sweep [wu-etal-2023-boundless-das] Applied as a 4096x4096 rotation across 7 layers, two causal models ("Left Boundary"; "Left and Right Boundary", two conjoined booleans) reach interchange-intervention accuracy at or above Alpaca's 85% task accuracy [wu-etal-2023-boundless-das] Alternative causal models fit much worse and a random-rotation control floors around 0.60 IIA [wu-etal-2023-boundless-das] The alignment is robust to unseen brackets (no drop), changed instruction wording (-1%) and added irrelevant context (-2%) [wu-etal-2023-boundless-das] A randomly-initialized LLaMA-7B control aligns no better than a most-frequent-label baseline (~66%), confirming the effect is not architectural; the paper does not restate DAS's caveat that orthogonal rotation assumes rather than discovers linear structure [wu-etal-2023-boundless-das]

Context

Boundless DAS (learned subspace dimensionality), interchange intervention accuracy at LLM scale, out-of-distribution / instruction-generalization robustness testing

Confirmed in models

Papers

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca — Wu, Zhengxuan, Geiger, Atticus, Icard, Thomas, Potts, Christopher, Goodman, Noah D.2023 · arXiv:2305.08809