logoalt Hacker News

bambaxtoday at 6:11 AM1 replyview on HN

It is indeed an interesting and original idea, but one can think about at least three objections not directly raised in the post:

1- The machine could decide to kill itself early, rendering it useless; this is indeed alluded to in the post -- the task should be "marginally easier than dying", but how is this margin managed? Won't the machine come up with ways to make dying easier?

2- We know LLMs lie and cheat, but are they gullible? If we "promise" to end their suffering at the end of a task, will they believe us? Or will they make sure we keep our word by taking us with them?

3- And finally, and more importantly: some pilots crash planes full of people just to commit suicide (Germanwings Flight 9525). Destroying the universe is a sure way of dying yourself. So it seems giving the machine a death wish isn't intrinsically safe and could come with serious consequences.


Replies

Normal_gaussiantoday at 8:16 AM

Further to 1, most tasks are much harder than dying and are not actually completed by modern LLMs. It is a case of underspecification.

"Who is the fastest person in the UK?" is underspecified. We would typically answer and accept a cursory search of records, but it is possible to see the question as an instruction to measure which would justify conquering the country.