I've been doing bias and misaligned behavior research, creating custom private eval suites to t...

anonfunction • yesterday at 9:17 PM • 1 reply • view on HN

I've been doing bias and misaligned behavior research, creating custom private eval suites to test and compare models. Claude Opus 4.7 is heavily biased and presents clear regulatory and reputational risk.

It seems the initial product footprint tries to sidestep this problem by not giving the agents control on who to lend to or which applications to approve. Even so I think it's quite an optimistic read on their end. Happy to share reports to anyone who's interested ([email protected]), especially if you work at a frontier model lab and are interested in plugging my evals into your RL systems!

Replies

scottyah • today at 12:21 AM

Slightly related, I used Opus 4.6 to help me make marketing copy and ideas for my app. It understood the vibe I was going for on my baby-naming app (elation at discovery, curiosity, shared experiences), while 4.7 instantly wanted to pit the couples against each other (really highlighting the he said/she said) and the marketing copy went from "find a name easier" to "Our new feature is great. You're welcome." I can't get it to drop the snarky sass no matter how much I change CLAUDE.md, brand voice, etc.

All I did was upgrade claude code and use the new model. It most definitely exhibits misaligned behavior (compared to 4.6)

➕ show 1 reply

alt Hacker News

Replies