logoalt Hacker News

mynti • today at 7:45 AM • 0 replies • view on HN

Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?