Unfortunately when it comes to language and intelligence the human population is exceptionally ignorant about it. I like Michael Levins work on intelligence at scale. We are kind of ok at seeing intelligence at human scale but it tends to fall apart after that. When you get to things like non-conscious intelligence humans are pretty bad at that too.
We could be (and are) encoding all kinds of behaviors in LLMs that are not at the word or token level. They are higher dimensional constructs. You won't see these things in the output of the prompt. A kind of subconscious (unstated in tokens) knowing that affects the output.