I think it’s just where they focused RLHF resources. The models have generally only gotten worse at writing.
And writing doesn’t have validators like code so you can’t really scale it in the same way