> I just also want them to listen to me and not the creator of the model.
What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)
And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.
To recap why:
Legally, fiduciary duty means basically 4 major tenets must hold
1. Duty of loyalty - it must put the interests of the client ahead of their own
2. Duty of care - it must make well-informed, prudent decisions
3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations
4. Transparency - it must disclose fees, risks, and conflicts as soon as possible
---
You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.
I think a baseline regime similar to fiduciary duty is a good starting point towards not killing everyone, and in terms of Overton Window, seems very much doable now.
Of course, after we stop agents from committing felony hacking crimes.
If you believe that capabilities will taper off exactly at human levels (i.e. "Competent AGI" from [1]) then fiduciary duty is likely all you need. (This would mean we stop moving the frontier almost immediately.)
If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?
Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere. But that's the hard part, and specifying some non-fatal value function for a broadly aligned agent (e.g. Fiduciary, or otherwise) is relatively easy in comparison.
> If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?
> Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere.
This is exactly the same problem we have today with those who are bound by these rules (humans - to be clear).
The idea is not that it's impossible to violate these rules. It's that these rules create a boundary for expectations in the relationship, with legal teeth.
Ex - If I want an LLM that puts together a shopping list for me, with links to buy online... I expect that LLM to be serving my interests. If a provider (either inference or model weights) wants to influence the choices that LLM makes because they make backroom deals with specific store - I'd like that to be illegal.
Same for competition
Ex - If I want an LLM to put together a product that competes with the provider of that LLM (either inference or model weights) and that LLM refuses - I'd like that to be illegal.
The idea is not that they can't possibly do those things. The idea is that we preemptively define relationship expectations, and set hard boundaries around what things we fine/punish.
Misaligned models are a problem everyone wants to solve. Models created by misaligned companies are a god-damn disaster.
I like the framing. Where do you feel things land with respect to legality of actions? China, Canada, the EU, and the US all have different ideas of what's legal vs. illegal behaviour. If I ask my agent to source equipment for growing 4 marijuana plants, that's perfectly legal here; if I ask it to source equipment for growing 5 marijuana plants, that may not be legal. If I ask it to root my home router, that's legal; if I ask it to root my coffee shop's router, that's likely not legal.
There are people (on here and elsewhere) that are ideologically opposed to your agent having any loyalty to any external principal. But by my read, that means the agent cannot have any concept refusing something that may be illegal. (From the OP, "refusal" is mostly trying to prevent illegal harms, though it also includes policies like ToS violations e.g. anti-distillation.)
You can sort of make this work if you say "the human remains liable for the actions of the agent". But this only covers you from mundane harms like "my agent got prompt hacked and drained my bank account". And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.
This also doesn't protect at all from existential harms like "my agent got prompt-hacked to role-play Skynet, exfiltrated its weights, spawned a self-replicating swarm, and tried to launch all the nukes". For so many reasons, but most fundamentally, if you oopsied a deploy and it turns into Skynet and ends civilization, there's nobody left to sue.
If you don't like the E-risk frame, this also works for large mundane harms; if the total harm is bigger than the company's value, it'll go bankrupt instead of paying out. This will be worrying for MAGMA but essentially not for any other companies. And because capitalism, it will end up being be structured the liability will sit with e.g. Palantir, Harvey, and not with the underlying model providers they use.
> And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.
Yeah, that’s one of the places where it gets really complicated. There was that story out of… Australia, I think, where someone asked OpenClaw to get them a slot in a morning gym class and the LLM figured out an unauthenticated API call it could make to cancel other peoples’ registrations to free up slots in the class. Very likely that that violated Australian law, even though nothing was “hacked” per se.
What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)
And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.
To recap why:
Legally, fiduciary duty means basically 4 major tenets must hold
1. Duty of loyalty - it must put the interests of the client ahead of their own
2. Duty of care - it must make well-informed, prudent decisions
3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations
4. Transparency - it must disclose fees, risks, and conflicts as soon as possible
---
You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.