in my experience it's very similar to what you typically need to do in order to get the best results in other fields (like audio), where being really particular/deliberate in how you manicure and format the input can pay orders-of-magnitude dividends
sometimes adding noise over a frequency range is better than removing it entirely, the opposite tends to be true for text and especially code, where you'll want lists of language and framework keywords, and then a pass on top of that to scan the codebase itself for its own slang. You can then 'double dip' and use these lists to weight the results after the fact