There are certainly uses where it’s good enough, but you can never be 100% certain of correctness in the way that people claim you can by stacking N layers of these models.
What triggered my response was the “just review the output with another LLM and it’s perfectly correct”