Linear accessibility is provably costlier than linear representation alone
measured in 1 paperGarg, Kleinberg & Peng split the linear representation hypothesis into linear representation (activations are Az) and linear accessibility (each feature is recoverable by a genuinely linear probe) [garg-etal-2026] For k-sparse features among m, embedding dimension d = O(k^2 log m) suffices for both, so superposition is achievable even under the stronger linear-accessibility requirement [garg-etal-2026] But d = Omega((k^2/log k) log(m/k)) is necessary, a provably larger (quadratic in k) requirement than classical compressed sensing needs with a nonlinear decoder [garg-etal-2026] Satisfying the capacity bound does not force orthogonal per-feature directions; the clean orthogonal picture returns only once representation and probe vectors have comparable bounded norms [garg-etal-2026]