If your output is based on best next result, it will tend toward the mode, not even the median.
I have to assume the most common [insert thing] is going to be bad, because more people are amateurs at [insert thing] than are experts. Right?
I have never expected excellence from chatbot generated anything. I have always expected just good enough.