logoalt Hacker News

ur-whaletoday at 3:13 PM3 repliesview on HN

What exactly prevents anyone worried about this to build an LLM that can decode the neuralese to english and use it to monitor what the model is doing?

Or is the neuralese some sort of irreversibly encrypted data set that only an LLM can "understand" and that can never be translated back to English?


Replies

LinchZhangtoday at 5:53 PM

It's very hard in practice. Even without optimization against it, it's about as hard as understanding an activation layer today and will probably get harder in the future.

There's also nothing stopping the secondary LLM from confabulating bullshit and there are less checks on it than CoT (harder for either humans or other models to externally verify).

Arodextoday at 3:17 PM

How do you trust that LLM?