logoalt Hacker News

user43928today at 4:04 PM0 repliesview on HN

About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.