A real EHR foundation model trained from scratch on MIMIC-IV recovers a correctly-ranked ordinal 1D continuum of decile embeddings via PCA
measured in 1 paperBurkhart, Ramadan, Liao, Chhikara, Rojas, Parker & Beaulieu-Jones (2025) train a real Llama-3.2-1B-architecture transformer from scratch on tokenized real EHR sequences (MIMIC-IV, CLIF-standardized), then transfer to a real independent hospital dataset (UCMC); two-component PCA of the model's own decile/quantile token embeddings recovers the correct ordinal ranking of all ten deciles as a 1D continuum, alongside category-consistent clustering of clinical-concept tokens [burkhart-etal-2025-foundation-models-for-electronic-health-records] A linear-probe (logistic regression) analysis of patient-representation trajectory features over the first 24 hours of admission (path length, max jump, anomaly score) predicts four clinical outcomes (ROC-AUC 0.877-0.914 in MIMIC, degrading to 0.529-0.878 in the independent hospital before fine-tuning) -- purely observational, no causal validation [burkhart-etal-2025-foundation-models-for-electronic-health-records]