Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
Have you considered it isn't?
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Meaning OP is a better programmer then a statistical model which produces the most probable results?
It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.
I’m saying it’s probably multiple factors and both you and GP are right.