Concept classes are multi-cluster; GMM-OT steering beats single directions
measured in 1 paper- Nominal concept classes used for activation steering (e.g. "harmful instructions") are internally heterogeneous: last-token hidden states form multiple semantically coherent sub-clusters (via k-means + PCA/t-SNE), not single Gaussian blobs. [abdullaev-etal-2026-concept-heterogeneity-representation-steering] - CHaRS steers by modeling source and target activations as Gaussian mixtures and applying an entropy-regularized Wasserstein optimal-transport map between cluster centroids (a barycentric combination of per-cluster diff-in-means shifts), degenerating to plain diff-in-means at one cluster. [abdullaev-etal-2026-concept-heterogeneity-representation-steering] - The steering-direction covariance is low-rank, rank(Sigma) <= 2K-2, motivating a spectral variant (CHaRS-PCT) that uses fewer directions. [abdullaev-etal-2026-concept-heterogeneity-representation-steering] - As a causal intervention (Activation Addition / Directional Ablation), CHaRS raises jailbreak attack-success rate up to 95.19% (Qwen2.5-7B-Instruct) over single-direction baselines, sweeping K=1 to 15 clusters. [abdullaev-etal-2026-concept-heterogeneity-representation-steering] - On RealToxicityPrompts it cuts toxicity up to 43% (CLS) and 42% (zero-shot) versus Linear-AcT while preserving perplexity/MMLU. [abdullaev-etal-2026-concept-heterogeneity-representation-steering] - Tested on seven instruct LLMs for jailbreak (Gemma2-9B, Llama3.1-8B, Llama3.2-3B, Qwen2.5-3B/7B/14B/32B), three for toxicity (Gemma2-2B, Llama3-8B, Qwen2.5-7B), plus FLUX.1 for image-style control (Mistral-7B as perplexity scorer). [abdullaev-etal-2026-concept-heterogeneity-representation-steering]