I suppose then the benchmark should be to give it numerous synthetic tax codes written with the same language and conventions as the target tax code and then have it file "tax returns" against the respective tax codes. The training to game such a benchmark then moves from "this specific tax code" to "tax codes like this"