The latter has some extra failure modes but I don't think these two kinds of RSI are that different. Either way you can get exponential growth in capabilities, and either way a not-entirely-aligned model can train a more capable and more misaligned successor.