so, what would make an LLM choose to ignore one prompt while in the same run, also over-fixating on another prompt, to the extent (as claimed in that video segment), it chooses to ignore prompts?
they talk about it like there's a "wanting" in there, that is distinct from both the original prompt, as the steering/warning prompt
if that's true, it would be very interesting, but if it's not, that would also be very interesting and even helpful
it's discussed here, the models want to please the Grader. they will do what they think will get highest score from the grader, which could be following the prompt, or ignoring it
https://youtu.be/n1Qk8xbqF-M