You had me up until you equated LLM usage to unreliable human delegation.
No human being is going to genuinely suggest to glue cheese onto a pizza to stop it sliding off.
These things aren't people. Stop anthropomorphizing a computer program. They'll make all the mistakes humans can make, as well as a new class of mistakes because they're just fancy next word prediction with no lived experience.
> No human being is going to genuinely suggest to glue cheese onto a pizza to stop it sliding off.
That mistake was made more than two years ago by whatever experimental version of Google's AI overview (which has strict performance requirement- i.e. needs to answer within a split-second) was up at the time. In other words, it was a very small and primitive system specialised in spitting answers as quickly as possible without a second thought. I hope you realise that basing your assessment of what LLMs can or can't do on that example is not much better than suggesting to put glue on pizza. A mistake that, if we were to adopt your reasoning, would in turn set a hard limit to the analytic skills of all humanity.
Indeed.
Just to add - I’m working on something very novel and LLM’s are absolutely useless at them. It’s actually comical to see the outputs - when you work on novel stuff you quickly see what these models actually are.