MATH · IN · MODELS
methods / Direction Extraction / Linear Artificial Tomography (LAT)

Linear Artificial Tomography (LAT)

Techniqueintermediate

Extracts a 'reading vector' for a high-level concept or behavioral function as the first principal component of paired activation differences, collected from contrastive stimulus templates designed to isolate that concept — a PCA-of-differences variant of direction extraction, usable with or without labels.

Used in (2 observations)

structure: Linear Direction · models: Llama-2-7B-Chat, Llama-2-13B-Chat, Llama-2-70B-Chat, Vicuna-13B, Vicuna-33B-Uncensored, DeBERTa-xxlarge-v2-MNLI · paper: Representation Engineering: A Top-Down Approach to AI Transparency
structure: Linear Direction · models: Llama-3-70B-Instruct, Gemma-7B-it, Qwen-1.8B-Chat, Vicuna-13B, Qwen1.5-1.8B-Chat, Qwen1.5-32B-Chat, Llama-2-13B-Chat, Llama-3.1-8B-Instruct, NeuralDaredevil-8B-abliterated, Hermes-2-Pro-Llama-3-8B, OLMo-7B-SFT, Zephyr-7B-Beta, H2O-Danube3-4B-Chat, Gemma-2-9B-it, Qwen2.5-7B-Instruct · paper: Refusal in Language Models Is Mediated by a Single Direction, Representation Engineering: A Top-Down Approach to AI Transparency, Programming Refusal with Conditional Activation Steering, Refusal Direction is Universal Across Safety-Aligned Languages