logoalt Hacker News

satvikpendemtoday at 5:14 PM3 repliesview on HN

Reduce or turn off thinking:

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates


Replies

me_bxtoday at 7:41 PM

In my experience with Qwen3.6 35B-A3B, disabling thinking made the model generate inaccurate replies. Ask it for the recipe of egg salad and it gives you the recipe of an omelette.

Did I miss something, is it possible to have that model be reliable without thinking?

Casteiltoday at 5:44 PM

Given that it apparently defaults to 'xhigh', this is probably the answer.

Granted, it's still much lower tokens/s than you'll get out of many MoE models.

Edit: Even set to medium or low there's still a lot of second guessing, less consistency, lower 'acceptable response' rate, and slower/more token churn vs gemma4:26b-a3b. I think gemma4 is just a better 'general purpose' model.

IronWolvetoday at 5:41 PM

Thank you, this is exactly what I needed.