logoalt Hacker News

thefxperson • today at 6:41 PM • 0 replies • view on HN

My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer.

Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:

  it remains unclear to what extent these performance gains can be attributed to human-like task decomposition or simply the greater computation that additional tokens allow. [...] our results show that additional tokens can provide computational benefits independent of token choice. The fact that intermediate tokens can act as filler tokens raises concerns about large language models engaging in unauditable, hidden computations that are increasingly detached from the observed chain-of-thought tokens.
https://arxiv.org/html/2404.15758v1