Well in this case that’s a valid complaint. The tax code changes yearly. If the primary reason an llm can get through a return is because it’s been trained on some example returns then it’ll work for a year. Which I don’t disagree that’s useful if you can retrain yearly but it doesn’t mean the problem has been solved naively.
I suppose then the benchmark should be to give it numerous synthetic tax codes written with the same language and conventions as the target tax code and then have it file "tax returns" against the respective tax codes. The training to game such a benchmark then moves from "this specific tax code" to "tax codes like this"
Tax codes are a great example of something changing frequently, and something very nuanced. So, the training data doesnt exist for a future year's return while are rules are already changed.
So, instead of math/programming, tax codes approximate the real world better. Whereas math/programming/images/languages are frozen in time. I find this fascinating. Almost like the legal system, wherein rules are laid out in english but are almost mathematical in the sense that there is a legal/illegal judgement made for almost every action, but that judgement is not easily determined.