logoalt Hacker News

sdoeringtoday at 11:04 AM1 replyview on HN

In my daily work, I found expeciall terra and sol now stopping every few rounds again, telling me the tak is done. I even had them create a detailled plan - and told them to finish "end to end" - and they appruptly stop after the plan. Because they interpret this as finished. Even if the DOD is clearly not "finish the plan".

The new models are shite (pardon my French), when it comes to long running tasks and I find myself more and more using open wheights models or switching back to gpt-5.5 for "real work".

This might be the fact, that i use them for non coding work. But the degradation between 5.5 and 5.6 is stark in my daily work.

As always with AI - everybody's mileage will vary.


Replies

toshtoday at 11:16 AM

Interesting, that is not my experience (but I'm mainly using them for reading and writing code right now)

but I don't doubt that you're seeing this behaviour, ty for sharing!