Custom Research Transformer (small, purpose-built for interpretability studies)
various (academic)
Structures found in this family (6)
By model (7)
Custom 3-layer decision-pretraining Transformer (Fang & Rajan 2026) · 3 layers, 512-dim embeddings, trained via decision-pretraining (DPT) meta-RL on gridworld and tree-maze tasks
Custom Transformer (trained from scratch on a structured synthetic knowledge graph) · small (research-scale)
Custom controlled 7B Transformer (trained from scratch for architecture/normalization ablations) · 7B
Graph Transformer (GT/GraphiT/SAN), trained on LRGB peptides-func/peptides-struct (Tori, Bini, Sorbi, Marchand-Maillet & Ginis)
TS-51M (custom 8-layer x 512d x 16-head transformer, TinyStories, 6 pretraining seeds) · 51M
Custom GPT-2-style causal transformer (205M, TinyStories, largest of a 3.4M-205M width/depth/epoch grid) · 205M
Custom ViT (10-layer, 8-head, 384-dim, trained from scratch with VICReg/SimCLR objectives on CIFAR-100/FOOD101) · 10-layer, 384-dim
no structures recorded for this checkpoint specifically
Observations (6)
Papers
From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers (2026), Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing (2024), Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026), The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks (2026), Probing Graph Neural Network Activation Patterns Through Graph Topology (2026), Linguistic Collapse: Neural Collapse in (Large) Language Models (2024)