logoalt Hacker News

pixl97yesterday at 6:55 PM1 replyview on HN

There are two questions about that stability I have.

One, things like catastrophic forgetting and falling into incoherence.

Two, less likely but far more worrying, falling into unwanted attractor states. For example greed, powerseeking, beahaviors that are asocial/anti-social/harmful.


Replies

fc417fc802today at 2:50 AM

Aren't such attractors also problems during training? Presumably alignment constraints would need to apply to continuous learning as well.

show 1 reply