If a change makes sense, a different frontier model can usually figure out why it was made from the code alone. I take that as a signal that that the commit is good.
I believe they can do this due to having millions of public pull requests and github issues in their training data.
There's pretty much always several plausible reasons for why a change was made. What's interesting is knowing exactly which one, especially when it later turns out to be wrong!
What a change does is self documenting. Why it was made and why that specific approach was chosen could be due to many things, such as:
* business objectives,
* the result of experimentation,
* the result of an offline conversation
etc.
This cannot be captured in code alone. This is why we write commit messages.