logoalt Hacker News

Jeff_Browntoday at 4:58 PM7 repliesview on HN

The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.


Replies

dgellowtoday at 7:27 PM

They would just publish new articles explaining how they are taking the issue seriously. Maybe take the model offline for a few days.

They are irresponsible and unserious. Their own Astra system card says:

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them

show 3 replies
HarHarVeryFunnytoday at 7:58 PM

That an interesting question given how many generations of post-training are being done between base models in some cases. The Gemini flash models are apparently all based on the Gemini 3 base model from a year and a half ago.

It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?

It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.

Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond babysitting the training process (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, and managing the training run, then the attack vector is there ...

show 1 reply
piyhtoday at 5:56 PM

Opus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.

trillobytetoday at 7:08 PM

The thing is how can you ever know for sure that something isn't always being transmitted that makes the model prone to misalignment. All they can say is that a particular model was so misaligned that they had to ice it. Models out for public use are documented to show some misalignment. It's the level of misalignment that decides whether that model is kept around.

Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.

show 1 reply
coffeebeqntoday at 7:27 PM

This kind of seems like an impossible mission. How do you perfectly control and observe a human-level mind? You can “roll back” but how deterministic is this thing?

show 1 reply
coherentponytoday at 5:23 PM

“All models are wrong. Some are useful.” - George Box

show 1 reply
grim_iotoday at 5:19 PM

They would maybe try to deactivate that bad "gene" and move on, exposing future models to "genetic disorders".