TCAV directional-derivative sensitivity testing
Techniqueintermediate
Given a concept activation vector (a linear-probe hyperplane normal), computes the directional derivative of a class logit along that direction for each input, then aggregates the sign of this derivative across an entire class into a single score (TCAVQ), with a statistical-significance test across many independently-trained CAVs to reject spurious concepts.
Used in (3 observations)
structure: Linear Direction · models: OpenFlamingo-3B-Instruct, OpenFlamingo-4B, OpenFlamingo-9B · paper: Visual concept ranking uncovers medical shortcuts used by large multimodal models
structure: Linear Direction · models: GoogLeNet (ImageNet-trained), Inception V3 (ImageNet-trained) · paper: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
structure: Linear Separability · models: Silverman 4x5 CNN (self-play chess RL agent), Los Alamos 6x6 ResNet-CNN (self-play chess RL agent) · paper: Reinforcement Learning in an Adaptable Chess Environment for Detecting Human-understandable Concepts