> We need independent audit and monitoring systems to assess the intent of each task and align it - in real time. This is far harder than it may first appear.
I may be too close to the research, but it appears to me to be so hard as to be unrealistic.
I recall some story a while back where an auditor wanted to see all TCP packets printed out on paper, and it had to be explained to them that this would require a continuous supply of trucks.
Tokens are regularly priced in cents or single digit dollars per million tokens. It's not quite a word per token, but yeah, nobody's reading all that.
Worse, we don't always know the intent even when looking. We have a few tools to attempt it, for example the (misleadingly named) "chain of thought", but that's more like a notepad and the better models get the more they can, for lack of better words, read (and write) between the lines. We have probes and J-space* is the most recent one I'm aware of, but we are still scratching the surface with how reliable and general these are.
But you said "need"; the need for something can be present without that thing being possible.
I agree on all points. It gets worse: OpenAI is switching their model thinking from sequential language tokens to primarily "latent neural representations." Meaning there will be little or no chain of thought to monitor. This appears more efficient, so all model labs will eventually switch to this. At the most crucial time for us to be monitoring reasoning and intent, we're about to make that much harder.
I also think the intent problem overlaps a frustrating amount with philosophical and political questions. It's the basis for Asimov's Three Laws of Robotics (1942). Intent is subjective. Language is subjective. Humans are imperfect at using language to accurately portray intent. All of these guarantee that an enormous number of queries in the future are going to be misinterpreted. Not such a big deal when it's about a cake recipe, but when it's about governance, laws, military targets, nuclear power sites, etc, the scope for failure becomes catastrophic. The Three Laws of Robotics attempt to create a backstop, but as countless stories have explored since (including I, Robot), even these laws are subject to interpretation.