Interesting take, given that the whole reason LLMs are interesting is that they're the first system we have that can work in semantic space. Statistical interpolation, that we've solved long ago.
I think that's equivocating on the word semantic. Word embeddings map concepts like "king - man + woman = queen" or cluster synonyms together in high-dimensional vector space and the ML literature very loosely calls this "semantic space." But a high-dimensional topology of token co-occurrences isn't semantics in the sense of computation or formal semantics. It's still "just" measuring distributional similarity. Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
Claiming we "solved" statistical interpolation long ago just means curve-fitting and basic regressions on structured data. Transformers are a truly impressive achievement, scaling all of this to unstructured high-dimensional text topologies, but it's fundamentally the same math operations on statistical proximity.
Like how do we explain hallucinations here? Tokens that are hallucinated are semantically "close" in that vector space but they're completely false in reality. If LLMs operated in a true semantic space they wouldn't hallucinate CLI flags that don't exist.
I think that's equivocating on the word semantic. Word embeddings map concepts like "king - man + woman = queen" or cluster synonyms together in high-dimensional vector space and the ML literature very loosely calls this "semantic space." But a high-dimensional topology of token co-occurrences isn't semantics in the sense of computation or formal semantics. It's still "just" measuring distributional similarity. Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
Claiming we "solved" statistical interpolation long ago just means curve-fitting and basic regressions on structured data. Transformers are a truly impressive achievement, scaling all of this to unstructured high-dimensional text topologies, but it's fundamentally the same math operations on statistical proximity.
Like how do we explain hallucinations here? Tokens that are hallucinated are semantically "close" in that vector space but they're completely false in reality. If LLMs operated in a true semantic space they wouldn't hallucinate CLI flags that don't exist.