How do you handle context limits? With more thinking tokens you fill it up earlier. Compaction degrades performance too. What's your strategy?