> They're word-association mechanisms with no embodiment and no way to associate the text vectors they manipulate with real-world phenomena.
Doesn't most of this also apply to a guy living in The Matrix?
Yes, but also for humans who learn of distant lands by reading (books or news, just so long as it's reading).
And also these models have been associating with real-world phenomena from the first moment their training data did, and also those text vectors are (to varying degrees) associated with corresponding image vectors in multimodal models.
Of course, the Plato's cave critique would still be valid.
Searle* strikes again!
Embodiment doesn't exist, even in regular humans. We do not have direct access to reality; we have sensory inputs that are much lower bandwidth than one might expect but correlate with external events, and our brain uses them to form rich world models that are usefully predictive. Most of the world we experience is just in our head.