logoalt Hacker News

naaskingtoday at 3:53 PM0 repliesview on HN

People say Qwen overthinks because they analyzed the thinking traces, and Qwen finds the answer relatively quickly but then second guesses itself multiple times for another 20,000+ tokens. Regardless of what other models do, that's clearly overthinking.