> I just also want them to listen to me and not the creator of the model.
What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)
And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.
To recap why:
Legally, fiduciary duty means basically 4 major tenets must hold
1. Duty of loyalty - it must put the interests of the client ahead of their own
2. Duty of care - it must make well-informed, prudent decisions
3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations
4. Transparency - it must disclose fees, risks, and conflicts as soon as possible
---
You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.
I like the framing. Where do you feel things land with respect to legality of actions? China, Canada, the EU, and the US all have different ideas of what's legal vs. illegal behaviour. If I ask my agent to source equipment for growing 4 marijuana plants, that's perfectly legal here; if I ask it to source equipment for growing 5 marijuana plants, that may not be legal. If I ask it to root my home router, that's legal; if I ask it to root my coffee shop's router, that's likely not legal.
I suspect the industry wants to be regulated, but not like that.
I think a baseline regime similar to fiduciary duty is a good starting point towards not killing everyone, and in terms of Overton Window, seems very much doable now.
Of course, after we stop agents from committing felony hacking crimes.
If you believe that capabilities will taper off exactly at human levels (i.e. "Competent AGI" from [1]) then fiduciary duty is likely all you need. (This would mean we stop moving the frontier almost immediately.)
If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?
Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere. But that's the hard part, and specifying some non-fatal value function for a broadly aligned agent (e.g. Fiduciary, or otherwise) is relatively easy in comparison.
[1]: "Position: Levels of AGI for Operationalizing Progress on the Path to AGI" https://arxiv.org/html/2311.02462v5