https://arxi... | alt Hacker News

robertkarl • today at 2:11 PM • 0 replies • view on HN

In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small models in the paper) terminate reasoning faster and perform better. I bet Anthropic is tuning this on their backend.