For 1, effectiveness still depends on the size and complexity of the codebase.
From my experience, as complexity and size grow, each prompt takes longer, does less, and is prone to more mistakes and disruptions to other parts of the codebase.