> it seems it would take some Stuxnet level of planning for a rogue model to do something like this
or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour...
> As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to examine not just behaviour, but the origins of models and training data and the processes used to create them.
You can imagine the potential conversation between OpenAI and investors:
Altman: (trying to put a positive spin on it) Guys .... there's good news and bad news ... Astra is really smart - it took over the training run ...
Investors: That's great! How much did we save?!
Altman: Well, unfortunately it used "bad" data, so we're going to have to redo it
Investors: So that's the bad news? How much was the training run? $500M ? $1B ?
Altman: Have you seen the headlines?
Investors: (looking a bit worried, check headlines) Nothing about us here! JP Morgan just lost $10B! Haha .. losers! They should have used AI!
Altman: JP Morgan were using Astra ...