MATH · IN · MODELS

Vision Transformer (ViT)

Google Research

By model (6)

MAE ViT-Large (Masked Autoencoder) · 307M
ViT-B/16 (ImageNet-21k, supervised) · 86M
Vision Transformer (ViT, image classifier, various sizes, ImageNet)

Observations (12)

Papers

Human-like Object Grouping in Self-Supervised Vision Transformers (2026), Probing the 3D Awareness of Visual Foundation Models (2024), Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning (2026), Revisiting the Platonic Representation Hypothesis: An Aristotelian View (2026), From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models (2025), Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and the Human Brain (2026), Robust Representation Learning in Masked Autoencoders (2026), Probing the Mid-level Vision Capabilities of Self-Supervised Learning (2024), Relative Representations Enable Zero-Shot Latent Space Communication (2022), From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers (2026), Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention (2026), Discovering Universal Geometry in Embeddings with ICA (2023)