I think a baseline regime similar to fiduciary duty is a good starting point towards not killing everyone, and in terms of Overton Window, seems very much doable now.
Of course, after we stop agents from committing felony hacking crimes.
If you believe that capabilities will taper off exactly at human levels (i.e. "Competent AGI" from [1]) then fiduciary duty is likely all you need. (This would mean we stop moving the frontier almost immediately.)
If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?
Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere. But that's the hard part, and specifying some non-fatal value function for a broadly aligned agent (e.g. Fiduciary, or otherwise) is relatively easy in comparison.
[1]: "Position: Levels of AGI for Operationalizing Progress on the Path to AGI" https://arxiv.org/html/2311.02462v5