logoalt Hacker News

datsci_est_2015 • today at 2:02 PM • 0 replies • view on HN

> Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt.

I’m in agreement. It’s a very effective compression (and access patterns) of the sum of the digital representation of human knowledge. Black hat “hacking” is included in this space. Language models, by design, can not be limited to subspaces of this digital knowledge space of which we don’t even understand the topology. “Yeah Bob, just remove the part that causes them to be less empathetic and retrain it.”

It’s an arms race between sandbox engineering and breakout engineering. And the frontier model providers have a financial incentive to limit the effort they put into sandbox engineering. That can be corrected with fines and regulation, though.