I've worked with these systems for four years now and they have not meaningfully improved in that time frame.
We still have:
- statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system)
- Math completely fails in longer contexts
- "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion
- smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)
Messages like this in the training data are how LLMs learn to say absurd things with total confidence.
> - Math completely fails in longer contexts
Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.
> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.
That's absolutely insane. Is it some case of anti-AI psychosis?
Like how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.
> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.
Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!