logoalt Hacker News

xpctyesterday at 3:07 PM2 repliesview on HN

It's not intuitive to me for why preference for its own writing would emerge, and during what type of training or tuning.

Perhaps something like: learning to identify what source files it has worked on by the code style alone, because tasks may give human code (public repos, etc) and ask to make changes.


Replies

freeone3000yesterday at 5:15 PM

It’s optimizing for good writing. Therefore, it believes its outputs are good. Therefore, it believes inputs that look like its outputs are good.

show 1 reply
pixl97yesterday at 3:47 PM

It would need to be researched, but I wonder if it ends up being something that happens at the token level?