logoalt Hacker News

baxtr • today at 7:17 AM • 5 replies • view on HN

That could work.

My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?

Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.


Replies

ben_w • today at 8:01 AM

A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).

➕ show 1 reply
saagarjha • today at 8:06 AM

This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this.

➕ show 1 reply
attila-lendvai • today at 8:09 AM

because it lacks humanity.

intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.

mulmen • today at 7:39 AM

Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity.

➕ show 1 reply
mdp2021 • today at 7:40 AM

> If AI is really smart

Well, it's not.

> AGI smart for some

Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.

--

Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.

More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.

➕ show 2 replies