logoalt Hacker News

Fordecyesterday at 9:48 PM5 repliesview on HN

This all strikes me as an effort to move tailoring the harness out of the easily transferable .md file into specific Anthropic tooling to increase lock in.

I've been running Opus 5 today and it's already done accidental deletions, made far more mistakes and worked around deliberate hook controls than previous Opus versions combined. Also it looks like token usage is up as it fails at the task the first time around much more frequently than 4.8.


Replies

frioyesterday at 10:18 PM

I’m not excited about using Opus 5, mainly because the way that I work atm — essentially peer programming — means I sandbox the agents and work with them closely. Opus 4.x encounters the sandbox and moves on with its day; Fable becomes increasingly fixated on it and does less and less of the actual task, focussing more and more on the limit it reached. I worry that, from your description, Opus 5 will do the same.

show 2 replies
vidarhyesterday at 10:27 PM

I have a document generation task that I used to run with 4.8. This morning after it switched to 5, the documents were consistently 30%-40% longer for the same prompt... Not evaluated whether they are actually better or worse yet, but what was interesting was how consistently more verbose it was.

show 1 reply
ValentineCyesterday at 11:41 PM

I haven't been impressed with Opus 5 over the past ~30 hours either.

It's made countless careless mistakes folding in plan amendments after they get reviewed by Sol, and has produced sloppy mockups (e.g. buttons overflowing past cards) despite all the supposed verification claims.

show 1 reply
slashdavetoday at 12:16 AM

Just make the contents of CLAUDE.md "@AGENTS.md"

Now, if only I could get Codex to read rules.

wren6991yesterday at 9:54 PM

I think the type of persistence rewarded by benchmarks may be misaligned with instruction following