MATH · IN · MODELS

NanoGPT Causal-Attention, No Positional Encoding (Causal-NoPE)

Research (Zuo, Guerzhoy & Guerzhoy 2025; custom-trained NanoGPT variants with causal attention and no positional encodings, on synthetic position-sensitive tasks)

Structures found in this family (1)

By model (1)

NanoGPT Causal-NoPE, 6-layer (10.6M params) · 10.6M

Papers

Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings (2025)