If you can do adversarial agent review, or use an LLM to fix the slop and it works, then technically you already have an LLM trained which knows what to do. You can just use them to improve the other model. The issue is that we don't have these at the moment.