A multiscale (macroscale identity-cloud, microscale intra-identity) geometric structure in face-recognition embedding spaces is measurably shaped by interpretable attributes, with FaceNet showing stronger attribute-dependence than ArcFace
measured in 1 paperLeroy, Mastropietro, Nurisso & Vaccarino (2025) identify a multiscale geometric structure in FaceNet and ArcFace embedding spaces: a macroscale organization of inter-identity point-cloud relations, and a microscale organization within each identity's own point cloud. Both scales are shown to be shaped by interpretable attributes (e.g. hair color, image contrast) via a physics-inspired "invariance energy" alignment metric applied to real pretrained face-recognition models and to fine-tuned variants trained with controlled attribute augmentation. FaceNet shows a higher structural dependency of inter-identity relations on attributes than ArcFace, and fine-tuning on a given attribute measurably increases that model's invariance-metric score specifically for the trained attribute. See [[attribute-induced-embedding-folding]].