Do you understand that LLMs are probabilistic?
Ask a model the same question twice and you will get different results. So, how were you ever getting “the best result, always”?