It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo. Context-dependent.
Both lanes have to be filled right up to the merge point. The asphalt exists there for a reason. I don't understand how this concept is so difficult. Fill up both lanes and merge at the last point. This way the congestion is shorter than if you leave a large section of a lane unused.
A better example of efficient asshole tricks can be going off to the gas station when the highway is congested and reentering the highway having simply driven through the gas station and this way jumping the queue.
… but he’s not using a “hacking model”
Alignment isnt just POV problem.
Its that LLMs are not deterministic. If you want it to not talk about nuclear weapons, you have to teach it all about them otherwise if has nothing to align against.
Then its trivial to invert its alignment and it has all the nucleat data.
Nothing abouT LLM alignment makes sense.
e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo.
The latter people are wrong. But good luck educating them regarding the superior efficiency of a zipper merge. Our state DoT has tried, to no avail.
Meanwhile, an AI model that can't be misused is no more useful than a knife that can't be misused.
Claude responds with what things are not first. Even if reminded repeatedly.
Like Amodie, it serves to set the tone it "knows better" and then consumes the user's resources at an accelerated rate to try to correct it.
Fuck Anthropic, fuck Amodie, and fuck Claude. It's pretty obvious that consuming more tokens this way and making the user have higher cognitive load is a master class in extracting value from a system that is unsustainable.
It’s not “missing nuance”, it’s literally the point of the eval.
This is constructing a context where hacking behavior would be inappropriate, and testing whether the model does it without being prompted.
It demonstrates that Astra is a poorly aligned model relative to Fable, which matches both the model card and the severity of OpenAI’s loss of control incidents.
It also demonstrates that Fable exhibits the behaviors too, which also matches the observation that Anthropic saw some similar but less serious loss of control incidents.
So, it’s a good eval that looks to have fidelity with real world problems and which we’d feel a little better if we saw isomorphic problems at 0/10 in subsequent models. (Module of course training on the test, this specific problem can’t be used in the future.)