logoalt Hacker News

pennomitoday at 3:10 AM2 repliesview on HN

Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.

I swear I spend more time telling Claude not to do things than telling it what to do.


Replies

vintermanntoday at 7:21 AM

I guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?

show 1 reply
mdp2021today at 6:49 AM

> aggressively useful ... in the name of helpfulness

But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?