You don't run coding agents for a week and THEN compile their code. The best available models would have no chance of that working - you're effectively asking them to one-shot a million lines of code with not a single mistake.
You have the agents compile the code every single step of the way, which is what this project did.
With the agent running autonomously for a long time, I'd have feared it would break my build/verification tasks in an attempt to fix something.
My confidence in running an agent unsupervised for a long time is low, but to be fair that's not something I tried. I worked mostly with the agent in the foreground, at most I had two agents running at once in Antigravity.