Is it really that simple ? A programming function that has set inputs and determinstic outputs is rather verifyable. A ten million lines code base is not easily verifyable. Knowing whether an agent is improving when it's working on a huge codebase is not that easy - it could easily start hurting code quality without it being clear in any output.