Models understand the relationships between words and outcomes, so the end result is the same. Whether they appreciate lie, cheat, and steal the same way as us is a philosophical question, not a practical one.
It is an important distinction - they do not 'understand' at all.
Input tokens map to output tokens. The illusion of comprehension is a byproduct.
Models don’t “understand” - they _encode_ the relationships between words.