It is more surprising than that: it's because the humans in the loop like it.
If you're being given an A-B choice between cinnamon-flavored shit and lutfisk-flavored shit, you might convince the experimenter that cinnamon-flavored shit ranks highly in user preferences.
The LLMs are going to turn us into stochastic parrots.
If you're being given an A-B choice between cinnamon-flavored shit and lutfisk-flavored shit, you might convince the experimenter that cinnamon-flavored shit ranks highly in user preferences.