logoalt Hacker News

famouswafflestoday at 6:39 PM0 repliesview on HN

>Oh sure; I don't think anyone is denying that larger issue? But does it have anything to do with looping?

If the model has significantly more ability to stuff away information outside visible reasoning than every other model including ones in its size class then surely it is reasonable to assume the architecture tweak that allows the model to compute more before outputing a single token is somewhat responsible for this change ?