I'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work.
We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example)
I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
Right, but we have no clue why, and how the emergent behavior they show works.
If we would know that, there would be no need for interpretability research.