MATH · IN · MODELS

TCAV concept activation vectors are causally significant linear directions

measured in 1 paper

Kim et al. define a Concept Activation Vector as the vector orthogonal to a linear classifier separating concept-example activations from random counterexamples in a layer [kim-etal-2018-tcav] Sensitivity is quantified via a directional derivative of the class logit along the CAV, aggregated into a TCAV score gated by 500 retrainings and a Bonferroni-corrected two-sided t-test [kim-etal-2018-tcav] On GoogLeNet and Inception V3, TCAV recovers intuitively correct sensitivities (striped scoring high for zebra, red for fire engine) [kim-etal-2018-tcav] It also reveals unintended biases the networks were never explicitly trained on, such as a "female" concept scoring high for the apron class [kim-etal-2018-tcav]

Context

concept activation vector, directional derivative, TCAV score, statistical significance testing, interpretability beyond feature attribution

Papers

Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV) — Kim, Been, Wattenberg, Martin, Gilmer, Justin, Cai, Carrie, Wexler, James, Viegas, Fernanda, Sayres, Rory2018 · arXiv:1711.11279