Cool project!
That's what I was insinuating through "better encoder"; the model creating more efficient representations of ASTs using something like JEPA