That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.
If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.
That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.
If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.