logoalt Hacker News

socializer • today at 5:14 AM • 0 replies • view on HN

I'll give you that the capabilities of LLMs are improving rapidly. Their safety, however, seems to be stuck in mid-2023. They're still easy to dupe and prone to cheating to solve problems. Prompt injection is still a thing, and it's still something we need to paper over with input and output classifiers and other hacks external to the LLM.

What we're seeing so far is consistent with the training data being the upper bound for capabilities. They get better at recall / synthesis / reasoning over the corpus, but they don't, for example, acquire trans-human ethics; they're at best as ethical as we are, except not grounded by the fear of consequences. A perfectly-behaved, perfectly-moral LLM is not a given in 20 years, not unless your position is that there's room for unbounded, exponential self-improvement without any loss of fidelity. In that case, we'll probably have problems more pressing than the outlook for accounting jobs.