I had a related insight, but in the domain of humor [1].
LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line).
In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context.
It might be possible to add this capability with a new architectural component like transformers, but specifically for making "left turns"/intuitive leaps.
You have committed a classic blunder of confusing your abstraction layers.
"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.
Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.
Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.
And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.
Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.
Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.