A cone probe isolates a causal tree-like subspace for genealogy
measured in 1 paperBaek et al. define a differentiable order-embedding cone probe (provably transitive and antisymmetric for a tree's descendant-of relation) and fit it on a 15-node family tree, then on 10-dimensional PCA-reduced activations of five instruction-tuned LLMs [baek-etal-2024-cone-probe-genealogical-representation-universality] Activation patching restricted to the fitted cone subspace produces a causal effect on genealogy Q&A comparable to or larger than a same-rank top-PCA subspace, but smaller than an unrestricted full-layer patch [baek-etal-2024-cone-probe-genealogical-representation-universality] The cone subspace thus captures real but incomplete causal structure, and a context-shuffle control degrades performance, consistent with reliance on depth-ordered structure [baek-etal-2024-cone-probe-genealogical-representation-universality] A complementary model-stitching experiment (OPT/Pythia/Mistral/Llama, 410M-8B) finds low next-token-loss degradation when splicing early/mid layers between independently-trained models [baek-etal-2024-cone-probe-genealogical-representation-universality] The authors explicitly hedge the universality conclusion given the lack of ground-truth graph representations [baek-etal-2024-cone-probe-genealogical-representation-universality]