1. Statistical Maximum Likelihood vs Factuality
Language models are trained via cross-entropy loss to predict the next token: $P(w_t \mid w_{1:t-1})$. They optimize for linguistic plausibility and corpus statistical frequency, not an immutable world-state knowledge graph. "Plausible sounding" sequences often dominate sparse factual truths.