Definition
Let be an embedding map and a one-parameter family of input transformations along attribute (e.g. = amount of hair-colour shift), with the identity. Define the attribute vector field
at each embedding point — the (normalized) direction in which moves under an infinitesimal push along attribute . Local invariance energy at over a neighbourhood radius is then
— explicitly analogous to the Hamiltonian of an -vector spin model, where plays the role of a local spin: low means the attribute-induced direction stays coherently aligned across a neighbourhood (an “ordered” local field), high means it varies erratically (a “disordered” one).
Intuition
A model made invariant to an attribute does not, in general, simply project that attribute’s information away or ignore it in a linear sense — it can instead achieve invariance by nonlinearly warping the space so that the attribute’s effect on the embedding is compressed or twisted rather than removed, “folding and distorting the space” (the paper’s own phrase) much like crumpling a sheet of paper contracts one direction more than others without deleting any point. The invariance energy quantifies exactly how coherent that fold is: a smooth, consistent fold across a neighbourhood gives low energy; an irregular one gives high energy, even for the same nominal degree of invariance measured by the model’s task accuracy.
Relative to manifold-curvature-profile
This is a directional, attribute-driven deformation field, not an intrinsic or extrinsic curvature scalar: Curvature profile of the representation manifold asks “how curved is the manifold at this point, independent of any labeled direction,” while attribute-induced folding asks “how does moving along one specific labeled attribute direction locally behave, and is that behaviour coherent across the space.” The two are compatible and could in principle be computed on the same embedding, but they are different quantities — a manifold could have uniform curvature everywhere while still folding erratically along a particular attribute’s direction, or vice versa.
Key evidence
A two-scale point-cloud structure in face-recognition embedding spaces — coarse inter-identity relations and fine intra-identity attribute-folding — is measurably shaped by interpretable attributes, with the degree of attribute-dependence differing between architectures. Leroy, Mastropietro, Nurisso & Vaccarino (2025) study FaceNet, ArcFace, and AdaFace face-embedding spaces at two explicitly distinct scales: macroscale (how attributes like hair colour or contrast relate different identities’ point clouds to each other, measured via a Kolmogorov-Smirnov statistic on intra- vs. inter-modality distance distributions — explicitly not discrete sub-clusters, unlike identity itself) and microscale (the attribute vector field and invariance energy defined above, measuring folding within a single identity’s point cloud). FaceNet shows systematically higher (steeper macroscale attribute-identity dependency) and significantly lower invariance energy across all attributes and scales (less coherent, more irregular folding) than ArcFace and AdaFace, and attribute-specific fine-tuning measurably raises the invariance-energy score for the fine-tuned attribute specifically. No formal named topological object (no stratified manifold, no fiber bundle) is claimed — the multiscale structure is an empirical, physics-inspired measurement framework built directly on point clouds, pairwise distances, and this tangent vector field. See face-recognition-embedding-spaces-show-multiscale-attribute-dependent-structure-differing-between-facenet-and-arcface.
Key papers
- Leroy, Mastropietro, Nurisso & Vaccarino (2025). Attributes Shape the Embedding Space of Face Recognition Models. arXiv:2507.11372 — origin of the attribute vector field and invariance-energy construction above.