BERT
Google
Structures found in this family (9)
Hypotheses argued for by this family (1)
By model (13)
BERT-base-uncased · 110.1M
BERT-base-cased · 110M
mBERT (BERT-base, Multilingual Cased) · 178M
BERT-large-uncased · 340M
BERT-large · 340M
BERT-large-cased · 340M
BERT-medium · 41.7M
BERT-mini · 11.3M
BERT-small · 29.1M
BERT-tiny · 4.4M
BERT base multilingual uncased · 168M
BERT-base-uncased fine-tuned on SNLI · 110.1M
BERT-base-uncased fine-tuned on SNLI + HELP · 110.1M
Observations (54)
Papers
Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction (2021), Interventional Probing in High Dimensions: An NLI Case Study (2023), How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings (2019), Investigating Aspect Features in Contextualized Embeddings with Semantic Scales and Distributional Similarity (2024), On the Sentence Embeddings from Pre-trained Language Models (2020), Visualizing and Measuring the Geometry of BERT (2019), Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models (2025), Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021), Evidence of Hierarchically-Complex Syntactic Structure Within BERT's Word Representations (2025), Finding Universal Grammatical Relations in Multilingual BERT (2020), Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations (2024), Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals (2021), DirectProbe: Studying Representations without Classifiers (2021), Emergence of Separable Manifolds in Deep Language Representations (2020), Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe (2026), A Closer Look at How Fine-tuning Changes BERT (2021), Mapping Semantic & Syntactic Relationships with Geometric Rotation in Embedding Space (2025), Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection (2020), Probing for the Usage of Grammatical Number (2022), IsoScore: Measuring the Uniformity of Embedding Space Utilization (2021), Editing Factual Knowledge in Language Models (2021), How Language-Neutral is Multilingual BERT? (2019), Language models align with brain regions that represent concepts across modalities (2025), Finding Lexical Identity and Inflectional Morphology in Modern Language Models (2025), Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes (2021), Cross-Lingual BERT Transformation for Zero-Shot Dependency Parsing (2019), Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color (2021), Emergent Linguistic Structure in Artificial Neural Networks Trained by Self-Supervision (2020), Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE (2025), Log-linear Guardedness and its Implications (2023), Naturalistic Causal Probing for Morpho-Syntax (2023), How Reliable are Causal Probing Interventions? (2025), The Representational Geometry of Number (2026), LEACE: Perfect Linear Concept Erasure in Closed Form (2023), Mapping Language Models to Grounded Conceptual Spaces (2022), How Multilingual is Multilingual BERT? (2019), A polar coordinate system represents syntax in large language models (2024), The Low-Dimensional Linear Geometry of Contextualized Word Representations (2021), Outlier Dimensions Encode Task-Specific Knowledge (2023), Linear Adversarial Concept Erasure (2022), The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024), All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality (2021), A Structural Probe for Finding Syntax in Word Representations (2019), Probing BERT in Hyperbolic Spaces (2021), Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings (2026), What are you sinking? A geometric approach on attention sink (2025), Transformer-Patcher: One Mistake Worth One Neuron (2023), On the Evolution of Syntactic Information Encoded by BERT's Contextualized Representations (2021), A Non-Linear Structural Probe (2021), WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach (2021), Whitening Sentence Representations for Better Semantics and Faster Retrieval (2021), Discovering Universal Geometry in Embeddings with ICA (2023)