logoalt Hacker News

absoluteunit1today at 7:54 PM2 repliesview on HN

> Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.

For some reason I had assumed testing this would be more sophisticated than just checking the thumbs up/down stats and user "vibes"


Replies

jonas21today at 8:33 PM

It's not just checking user thumbs up/down. As your quote says, they also did a controlled study with people rating the results. What else would you want them to do? The whole point is that it needs to introduce a detectable statistical difference, but humans should not be able to perceive it as a quality difference.

show 1 reply
cube00today at 8:03 PM

More unannounced testing on paying customers.

show 2 replies