logoalt Hacker News

javchztoday at 1:32 AM1 replyview on HN

I wonder if this can be fixed with LORAs.


Replies

bitexplodertoday at 1:41 AM

I had to fix this on 35B A3B -- I have a proxy that just shuts it down if it gets to 2K thinking tokens and injects something like "We have thought enough, let's begin working." and it almost always finishes the turn then. It rarely needs more than 2K thinking tokens and if it does there is always next turn. I would need to see what 27B is actually doing, but these smaller Qwen models seem prone to this.

show 2 replies