logoalt Hacker News

edottoday at 11:41 AM1 replyview on HN

But is that data good? That's the question. As in, is my usage at work:

a) indicative of problems that aren't already out there in the wild? (no) b) are the responses I'm getting so good and novel that the model can improve itself? (no)

It's the garbage in garbage out idea, just scaled up. If the model gave a bad answer, and I didn't catch it, and you now train on that I/O pair (my perhaps crappy prompt, the bad output), then you're not going to improve anything.


Replies

idiotsecanttoday at 1:22 PM

It seems like the user response rating mechanism might be a valuable signal